Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
✅ Single Commit Policy - COMPLIANTStatus: Policy requirements met • 1 commit • Valid format • Ready for merge 📊 View validation details📝 Commit Details
✅ Validation Results
🤖 Automated validation by NeuroLink Single Commit Enforcement |
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
|
Warning Rate limit exceeded
To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing. ⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (56)
WalkthroughAdded comprehensive speech-to-text, realtime voice, and new text-to-speech provider support. Introduced six STT providers (AssemblyAI, Azure, Deepgram, Gladia, Google, Whisper), two realtime providers (OpenAI, Gemini), four TTS providers, audio utilities for format detection and streaming, processor-based orchestration layers, and integrated STT into core generation workflows. Moved AWS SageMaker and Picovoice Cobra dependencies to required installation. Changes
Estimated code review effort🎯 5 (Critical) | ⏱️ ~120 minutes Rationale: Heterogeneous changes across 50+ files with high-density logic in provider implementations (6 STT adapters + 4 STT handlers + 4 TTS handlers + 2 realtime handlers = 16 complex providers), 500–700 LOC per provider. Core generation flow modifications, processor-based orchestration patterns, WebSocket/async streaming integration, comprehensive type system expansion (3 new type modules with 1900+ LOC), error hierarchy with factory patterns, and observability instrumentation. Diverse patterns (REST HTTP, WebSocket, async generators, event emitters) require separate reasoning per provider and module. Dependency reshuffling and CLI integration add complexity. Possibly related PRs
Suggested labels
Suggested reviewers
Poem
✨ Finishing Touches🧪 Generate unit tests (beta)
|
There was a problem hiding this comment.
Pull request overview
Adds a new voice/speech subsystem to the NeuroLink SDK + CLI, covering TTS, STT, and realtime voice providers, with supporting registries/factories and a continuous test script.
Changes:
- Introduces voice infrastructure (registry/factory, composite orchestration, realtime base processor, stream utilities, error types).
- Adds provider implementations for TTS (Google/OpenAI/ElevenLabs/Azure), STT (Whisper/OpenAI, Deepgram, Google, Azure), and realtime (OpenAI Realtime, Gemini Live).
- Extends the CLI with
neurolink voicesubcommands and adds a continuous voice integration test script.
Reviewed changes
Copilot reviewed 33 out of 33 changed files in this pull request and generated 8 comments.
Show a summary per file
| File | Description |
|---|---|
| test/continuous-test-suite-voice.ts | Adds a runnable continuous “smoke suite” for provider/module presence and basic wiring. |
| src/lib/voice/voiceRegistry.ts | Implements provider registry (IDs, aliases, metadata, type filtering). |
| src/lib/voice/voiceAgent.ts | Adds a high-level voice-to-voice agent (STT → LLM → TTS) + realtime session support. |
| src/lib/voice/stream-handler.ts | Adds chunking/backpressure helpers and async-iterable adapters for audio streams. |
| src/lib/voice/providers/OpenAITTS.ts | Implements OpenAI TTS via REST API. |
| src/lib/voice/providers/OpenAISTT.ts | Implements OpenAI Whisper STT via multipart/form-data REST API. |
| src/lib/voice/providers/OpenAIRealtime.ts | Implements OpenAI Realtime bidirectional voice over WebSocket. |
| src/lib/voice/providers/GoogleTTS.ts | Implements Google Cloud TTS via REST API + voice listing cache. |
| src/lib/voice/providers/GoogleSTT.ts | Implements Google Cloud STT via REST API + placeholder streaming. |
| src/lib/voice/providers/GeminiLive.ts | Implements Gemini Live bidirectional voice over WebSocket. |
| src/lib/voice/providers/ElevenLabsTTS.ts | Implements ElevenLabs TTS via REST API + voice listing cache. |
| src/lib/voice/providers/DeepgramSTT.ts | Implements Deepgram STT (batch + WebSocket streaming). |
| src/lib/voice/providers/AzureTTS.ts | Implements Azure TTS via REST API + voice listing cache. |
| src/lib/voice/providers/AzureSTT.ts | Implements Azure STT via REST API + placeholder streaming. |
| src/lib/voice/index.ts | Exports the voice module public surface area. |
| src/lib/voice/errors.ts | Adds shared error classes/codes for voice, STT, realtime. |
| src/lib/voice/compositeVoice.ts | Adds a TTS+STT orchestrator with history and conversation-turn helper. |
| src/lib/voice/audio-utils.ts | Adds audio format detection, duration estimation, and WAV/PCM utilities. |
| src/lib/voice/STTProvider.ts | Adds a centralized STT processor (similar to existing TTSProcessor pattern). |
| src/lib/voice/RealtimeVoiceAPI.ts | Adds centralized realtime processor + base handler abstraction. |
| src/lib/types/tts.ts | Expands AudioFormat union to cover more formats used by STT APIs. |
| src/lib/types/index.ts | Exports the consolidated voice types. |
| src/lib/adapters/stt/whisperSTTHandler.ts | Adds an adapter-style Whisper STT handler (non-voice-module path). |
| src/cli/parser.ts | Wires the new voice command group into the CLI. |
| src/cli/commands/voice.ts | Implements `neurolink voice synthesize |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| isConfigured(): boolean { | ||
| return this.apiKey !== null || this.credentialsPath !== null; | ||
| } | ||
|
|
||
| async getVoices(languageCode?: string): Promise<TTSVoice[]> { | ||
| if (!this.isConfigured()) { | ||
| throw new TTSError({ | ||
| code: TTS_ERROR_CODES.PROVIDER_NOT_CONFIGURED, | ||
| message: "Google TTS is not configured", | ||
| category: ErrorCategory.CONFIGURATION, | ||
| severity: ErrorSeverity.HIGH, | ||
| retriable: false, | ||
| }); | ||
| } | ||
|
|
||
| // Return cached voices if valid and no language filter | ||
| if ( | ||
| this.voicesCache && | ||
| Date.now() - this.voicesCache.timestamp < GoogleTTS.CACHE_TTL_MS && | ||
| !languageCode | ||
| ) { | ||
| return this.voicesCache.voices; | ||
| } | ||
|
|
||
| try { | ||
| const params = new URLSearchParams(); | ||
| if (languageCode) { | ||
| params.set("languageCode", languageCode); | ||
| } | ||
|
|
||
| const url = this.apiKey | ||
| ? `${this.baseUrl}/voices?key=${this.apiKey}&${params.toString()}` | ||
| : `${this.baseUrl}/voices?${params.toString()}`; | ||
|
|
||
| const response = await fetch(url, { | ||
| method: "GET", | ||
| headers: { | ||
| ...(this.credentialsPath && !this.apiKey | ||
| ? { Authorization: `Bearer ${await this.getAccessToken()}` } | ||
| : {}), | ||
| }, | ||
| }); | ||
|
|
||
| if (!response.ok) { | ||
| throw new Error(`HTTP ${response.status}`); | ||
| } | ||
|
|
||
| const data = (await response.json()) as GoogleListVoicesResponse; | ||
|
|
||
| const voices: TTSVoice[] = data.voices.map((voice) => ({ | ||
| id: voice.name, | ||
| name: voice.name, | ||
| languageCode: voice.languageCodes[0] ?? "en-US", | ||
| languageCodes: voice.languageCodes, | ||
| gender: this.mapGender(voice.ssmlGender), | ||
| type: this.extractVoiceType(voice.name), | ||
| naturalSampleRateHertz: voice.naturalSampleRateHertz, | ||
| })); | ||
|
|
||
| // Cache if no language filter | ||
| if (!languageCode) { | ||
| this.voicesCache = { voices, timestamp: Date.now() }; | ||
| } | ||
|
|
||
| return voices; | ||
| } catch (err: unknown) { | ||
| const errorMessage = | ||
| err instanceof Error ? err.message : String(err || "Unknown error"); | ||
| logger.error(`[GoogleTTSHandler] Failed to get voices: ${errorMessage}`); | ||
| throw new TTSError({ | ||
| code: TTS_ERROR_CODES.SYNTHESIS_FAILED, | ||
| message: `Failed to get voices: ${errorMessage}`, | ||
| category: ErrorCategory.NETWORK, | ||
| severity: ErrorSeverity.MEDIUM, | ||
| retriable: true, | ||
| originalError: err instanceof Error ? err : undefined, | ||
| }); | ||
| } | ||
| } | ||
|
|
||
| async synthesize(text: string, options: TTSOptions = {}): Promise<TTSResult> { | ||
| if (!this.isConfigured()) { | ||
| throw new TTSError({ | ||
| code: TTS_ERROR_CODES.PROVIDER_NOT_CONFIGURED, | ||
| message: "Google TTS is not configured", | ||
| category: ErrorCategory.CONFIGURATION, | ||
| severity: ErrorSeverity.HIGH, | ||
| retriable: false, | ||
| }); | ||
| } | ||
|
|
||
| const startTime = Date.now(); | ||
| const googleOptions = options as GoogleTTSOptions; | ||
|
|
||
| try { | ||
| // Detect if text is SSML | ||
| const isSSML = text.trim().startsWith("<speak"); | ||
|
|
||
| // Build synthesis input | ||
| const input: GoogleSynthesisInput = isSSML ? { ssml: text } : { text }; | ||
|
|
||
| // Parse voice and language from voice name or use defaults | ||
| const voiceName = options.voice ?? "en-US-Neural2-C"; | ||
| const languageCode = this.extractLanguageCode(voiceName); | ||
|
|
||
| // Build voice selection | ||
| const voice: GoogleVoiceSelectionParams = { | ||
| languageCode, | ||
| name: voiceName, | ||
| }; | ||
|
|
||
| // Build audio config | ||
| const audioConfig: GoogleAudioConfig = { | ||
| audioEncoding: this.getEncoding(options.format ?? "mp3"), | ||
| speakingRate: options.speed ?? 1.0, | ||
| pitch: options.pitch ?? 0.0, | ||
| volumeGainDb: options.volumeGainDb ?? 0.0, | ||
| }; | ||
|
|
||
| if (googleOptions.sampleRateHertz) { | ||
| audioConfig.sampleRateHertz = googleOptions.sampleRateHertz; | ||
| } | ||
|
|
||
| if (googleOptions.effectsProfileId) { | ||
| audioConfig.effectsProfileId = googleOptions.effectsProfileId; | ||
| } | ||
|
|
||
| // Build request | ||
| const request: GoogleSynthesizeRequest = { | ||
| input, | ||
| voice, | ||
| audioConfig, | ||
| }; | ||
|
|
||
| const url = this.apiKey | ||
| ? `${this.baseUrl}/text:synthesize?key=${this.apiKey}` | ||
| : `${this.baseUrl}/text:synthesize`; | ||
|
|
||
| const response = await fetch(url, { | ||
| method: "POST", | ||
| headers: { | ||
| "Content-Type": "application/json", | ||
| ...(this.credentialsPath && !this.apiKey | ||
| ? { Authorization: `Bearer ${await this.getAccessToken()}` } | ||
| : {}), | ||
| }, | ||
| body: JSON.stringify(request), | ||
| }); | ||
|
|
||
| if (!response.ok) { | ||
| const errorData = await response | ||
| .json() | ||
| .catch(() => Object.create(null) as Record<string, unknown>); | ||
| const errorMessage = | ||
| (errorData as { error?: { message?: string } }).error?.message || | ||
| `HTTP ${response.status}`; | ||
| throw new Error(errorMessage); | ||
| } | ||
|
|
||
| const data = (await response.json()) as GoogleSynthesizeResponse; | ||
| const latency = Date.now() - startTime; | ||
|
|
||
| // Decode base64 audio | ||
| const audioBuffer = Buffer.from(data.audioContent, "base64"); | ||
|
|
||
| const result: TTSResult = { | ||
| buffer: audioBuffer, | ||
| format: options.format ?? "mp3", | ||
| size: audioBuffer.length, | ||
| voice: voiceName, | ||
| sampleRate: | ||
| googleOptions.sampleRateHertz ?? | ||
| this.getDefaultSampleRate(options.format), | ||
| metadata: { | ||
| latency, | ||
| provider: "google-tts", | ||
| encoding: audioConfig.audioEncoding, | ||
| }, | ||
| }; | ||
|
|
||
| logger.info( | ||
| `[GoogleTTSHandler] Synthesized ${audioBuffer.length} bytes in ${latency}ms`, | ||
| ); | ||
|
|
||
| return result; | ||
| } catch (err: unknown) { | ||
| if (err instanceof TTSError) { | ||
| throw err; | ||
| } | ||
|
|
||
| const errorMessage = | ||
| err instanceof Error ? err.message : String(err || "Unknown error"); | ||
| logger.error(`[GoogleTTSHandler] Synthesis failed: ${errorMessage}`); | ||
| throw new TTSError({ | ||
| code: TTS_ERROR_CODES.SYNTHESIS_FAILED, | ||
| message: `Synthesis failed: ${errorMessage}`, | ||
| category: ErrorCategory.EXECUTION, | ||
| severity: ErrorSeverity.HIGH, | ||
| retriable: true, | ||
| context: { textLength: text.length }, | ||
| originalError: err instanceof Error ? err : undefined, | ||
| }); | ||
| } | ||
| } | ||
|
|
||
| /** | ||
| * Map SSML gender to standard gender type | ||
| */ | ||
| private mapGender(ssmlGender: string): "male" | "female" | "neutral" { | ||
| switch (ssmlGender?.toUpperCase()) { | ||
| case "MALE": | ||
| return "male"; | ||
| case "FEMALE": | ||
| return "female"; | ||
| default: | ||
| return "neutral"; | ||
| } | ||
| } | ||
|
|
||
| /** | ||
| * Extract voice type from voice name | ||
| */ | ||
| private extractVoiceType( | ||
| name: string, | ||
| ): "standard" | "wavenet" | "neural" | "chirp" | "unknown" { | ||
| const nameLower = name.toLowerCase(); | ||
| if (nameLower.includes("neural2") || nameLower.includes("neural")) { | ||
| return "neural"; | ||
| } | ||
| if (nameLower.includes("wavenet")) { | ||
| return "wavenet"; | ||
| } | ||
| if (nameLower.includes("standard")) { | ||
| return "standard"; | ||
| } | ||
| if (nameLower.includes("chirp")) { | ||
| return "chirp"; | ||
| } | ||
| return "unknown"; | ||
| } | ||
|
|
||
| /** | ||
| * Extract language code from voice name | ||
| */ | ||
| private extractLanguageCode(voiceName: string): string { | ||
| // Voice names are formatted like "en-US-Neural2-C" | ||
| const match = voiceName.match(/^([a-z]{2}-[A-Z]{2})/); | ||
| return match ? match[1] : "en-US"; | ||
| } | ||
|
|
||
| /** | ||
| * Get encoding string for audio format | ||
| */ | ||
| private getEncoding(format: AudioFormat): string { | ||
| const encodings: Partial<Record<AudioFormat, string>> = { | ||
| mp3: "MP3", | ||
| wav: "LINEAR16", | ||
| ogg: "OGG_OPUS", | ||
| opus: "OGG_OPUS", | ||
| }; | ||
| return encodings[format] ?? "MP3"; | ||
| } | ||
|
|
||
| /** | ||
| * Get default sample rate for format | ||
| */ | ||
| private getDefaultSampleRate(format?: AudioFormat): number { | ||
| switch (format) { | ||
| case "wav": | ||
| return 16000; | ||
| case "ogg": | ||
| case "opus": | ||
| return 24000; | ||
| default: | ||
| return 24000; | ||
| } | ||
| } | ||
|
|
||
| /** | ||
| * Get access token from service account (placeholder) | ||
| */ | ||
| private async getAccessToken(): Promise<string> { | ||
| logger.warn( | ||
| "[GoogleTTSHandler] Service account auth not implemented, use API key", | ||
| ); | ||
| return ""; | ||
| } |
There was a problem hiding this comment.
isConfigured() returns true when GOOGLE_APPLICATION_CREDENTIALS is set, but getAccessToken() is a placeholder that returns an empty string. In that configuration the provider will attempt authenticated requests with Authorization: Bearer and fail at runtime. Either implement service-account token acquisition (e.g., via google-auth-library) or treat "credentialsPath only" as not configured until token support is implemented.
| JSON.stringify({ | ||
| type: "conversation.item.create", | ||
| item: { | ||
| type: "function_call_output", | ||
| call_id: name, // Note: This should be the actual call_id from the event | ||
| output: JSON.stringify(result), | ||
| }, |
There was a problem hiding this comment.
call_id is being set to the function name, not the model-provided call identifier. OpenAI Realtime expects the original call_id from the function-call event; using the name will cause tool outputs to be ignored or misrouted. Capture call_id from the incoming event (e.g., include it in the parsed event shape) and echo that value here.
| const sessionConfig: RealtimeConfig = { | ||
| provider: "openai", | ||
| voice: this.config.voiceSettings?.voiceId, | ||
| instructions: this.config.systemPrompt, | ||
| temperature: 0.7, | ||
| turnDetection: "server_vad", | ||
| ...this.config.realtimeConfig, | ||
| ...config, | ||
| }; |
There was a problem hiding this comment.
RealtimeConfig consumers in this PR use config.systemPrompt (e.g., the provider session setup), but startRealtimeSession() sets instructions instead. This means the system prompt likely won’t be applied for realtime sessions unless providers also read instructions. Set systemPrompt (or set both fields consistently) when building sessionConfig.
| // Convert to voice module STT options format | ||
| const voiceSTTOptions: STTOptions = { | ||
| language: options.language, | ||
| format: options.format as STTOptions["format"], | ||
| diarization: options.diarization, | ||
| punctuate: options.punctuate, | ||
| wordTimestamps: options.wordTimestamps, | ||
| confidenceThreshold: (options as { confidenceThreshold?: number }) | ||
| .confidenceThreshold, | ||
| }; | ||
|
|
||
| const mergedOptions: STTOptions = { | ||
| ...this.config.defaultSTTOptions, | ||
| ...voiceSTTOptions, |
There was a problem hiding this comment.
transcribe() rebuilds a new STTOptions object with only a small subset of fields, which will silently drop valid options like sampleRate, punctuation, speakerDiarization, profanityFilter, etc. Since options is already Partial<STTOptions>, merge it directly with defaultSTTOptions (and only normalize aliases if needed) to avoid losing settings.
| // Convert to voice module STT options format | |
| const voiceSTTOptions: STTOptions = { | |
| language: options.language, | |
| format: options.format as STTOptions["format"], | |
| diarization: options.diarization, | |
| punctuate: options.punctuate, | |
| wordTimestamps: options.wordTimestamps, | |
| confidenceThreshold: (options as { confidenceThreshold?: number }) | |
| .confidenceThreshold, | |
| }; | |
| const mergedOptions: STTOptions = { | |
| ...this.config.defaultSTTOptions, | |
| ...voiceSTTOptions, | |
| // Merge caller-provided STT options directly so valid fields are not dropped. | |
| // Preserve the existing normalization for fields that require explicit typing. | |
| const mergedOptions: STTOptions = { | |
| ...this.config.defaultSTTOptions, | |
| ...options, | |
| format: | |
| (options.format as STTOptions["format"] | undefined) ?? | |
| this.config.defaultSTTOptions?.format, | |
| confidenceThreshold: | |
| (options as { confidenceThreshold?: number }).confidenceThreshold ?? | |
| this.config.defaultSTTOptions?.confidenceThreshold, |
| end(): void { | ||
| if (this.isEnded) { | ||
| return; | ||
| } | ||
|
|
||
| this.isEnded = true; | ||
| this.clearBufferTimeout(); | ||
|
|
||
| // Emit final chunk with remaining data | ||
| if (this.buffer.length > 0) { | ||
| const durationMs = | ||
| (this.buffer.length / | ||
| this.config.bytesPerSample / | ||
| this.config.sampleRate) * | ||
| 1000; | ||
|
|
||
| const chunk: AudioStreamChunk = { | ||
| data: this.buffer, | ||
| index: this.chunkIndex++, | ||
| isFinal: true, | ||
| format: this.config.format, | ||
| sampleRate: this.config.sampleRate, | ||
| timestampMs: this.timestampMs, | ||
| durationMs, | ||
| }; | ||
|
|
There was a problem hiding this comment.
If backpressure occurred, data is buffered in pendingData, but end() only flushes this.buffer and ignores pendingData. That can drop audio data if the caller ends the stream while paused. Consider concatenating/processing pendingData before emitting the final chunk (or ensure end() drains pending data first).
| // Detect format from extension | ||
| const ext = path.extname(inputPath).toLowerCase().slice(1); | ||
| const format = argv.format ?? ext; | ||
|
|
||
| spinner.text = `Transcribing with ${providerName}...`; | ||
|
|
||
| const options: STTOptions = { | ||
| language: argv.language, | ||
| format: format as STTOptions["format"], | ||
| diarization: argv.diarization, | ||
| wordTimestamps: argv.wordTimestamps, | ||
| }; |
There was a problem hiding this comment.
The CLI accepts/auto-detects formats like flac, m4a, webm, etc., but there’s no validation against the selected STT provider’s getSupportedFormats(). This can lead to confusing runtime failures (or wrong MIME types) when a provider doesn’t support the detected format. Validate the resolved format against provider.getSupportedFormats() and fail fast with a helpful message listing supported formats.
| const realtimeProviders = [ | ||
| { name: "OpenAIRealtime", envKey: "OPENAI_API_KEY" }, | ||
| { name: "GeminiLive", envKey: "GOOGLE_AI_STUDIO_API_KEY" }, | ||
| ]; |
There was a problem hiding this comment.
GeminiLive reads its API key from process.env.GOOGLE_API_KEY, but this test suite checks GOOGLE_AI_STUDIO_API_KEY. That mismatch will incorrectly report missing credentials (or skip unexpectedly). Update the env var used here to match the provider implementation (or support both).
| import * as fs from "fs"; | ||
| import * as path from "path"; | ||
|
|
There was a problem hiding this comment.
fs and path are imported but never used in this script, which will trigger lint/tsc unused import warnings. Remove them or use them for any planned file IO to keep the test runner clean.
There was a problem hiding this comment.
Actionable comments posted: 2
Note
Due to the large number of review comments, Critical severity comments were prioritized as inline comments.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/lib/neurolink.ts (1)
11666-11782:⚠️ Potential issue | 🟠 MajorThe lazy-init path creates an unusable
CompositeVoice.On a fresh
NeuroLinkinstance these methods callinitializeVoice()with{}, which creates aCompositeVoicewith nottsProvider/sttProvider. The very next call then fails insideCompositeVoice.synthesize()/CompositeVoice.transcribe()with provider-not-configured.🔧 Proposed fix
async synthesize(text: string, options?: TTSOptions): Promise<TTSResult> { - // Initialize voice if not already done - if (!this.compositeVoice) { - await this.initializeVoice(); - } - if (!this.compositeVoice) { throw new VoiceError({ code: VOICE_ERROR_CODES.INVALID_CONFIGURATION, message: "Voice not initialized. Call initializeVoice() first.", @@ async transcribe( audio: Buffer | ArrayBuffer, options?: STTOptions, ): Promise<STTResult> { - // Initialize voice if not already done - if (!this.compositeVoice) { - await this.initializeVoice(); - } - if (!this.compositeVoice) { throw new VoiceError({ code: VOICE_ERROR_CODES.INVALID_CONFIGURATION, message: "Voice not initialized. Call initializeVoice() first.",Also applies to: 11859-11878
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/neurolink.ts` around lines 11666 - 11782, The lazy-init path can create a CompositeVoice with no providers, causing CompositeVoice.synthesize()/transcribe() to fail; fix by adding explicit provider checks after initialization: in synthesize (and similarly in transcribe) verify this.compositeVoice and that a TTS/STT provider is configured (e.g., this.compositeVoice.ttsProvider / .sttProvider or a hasTTS()/hasSTT() helper) and throw a descriptive VoiceError (use VOICE_ERROR_CODES.INVALID_CONFIGURATION) that instructs callers to call initializeVoice(...) with a ttsProvider/sttProvider or configure SDK defaults; alternatively, update initializeVoice to pull default providers from NeuroLink configuration when config is empty so CompositeVoice is created with usable providers. Ensure references: initializeVoice, synthesize, transcribe, CompositeVoice, VOICE_ERROR_CODES, VoiceError.
🟠 Major comments (25)
test/continuous-test-suite-voice.ts-303-310 (1)
303-310:⚠️ Potential issue | 🟠 MajorMake missing
audio-utilsexports fail this suite.
exists || trueforces the check to pass every time, so a missing export only shows up as skipped noise instead of a real failure. This turns the export verification into a no-op.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@test/continuous-test-suite-voice.ts` around lines 303 - 310, The test loop is forcing every export check to pass by using `exists || true`; change the pass condition to use the actual `exists` boolean so missing exports fail the suite. Locate the loop that iterates `for (const fn of expectedFunctions)` and update the `recordTest` call for `audio-utils.${fn} exists` so the second argument is `exists` (not `exists || true`) and keep the failure message (`exists ? undefined : "Function not found"`) as-is; this ensures `audioUtils` missing exports cause a real test failure.src/lib/voice/providers/AzureTTS.ts-164-196 (1)
164-196:⚠️ Potential issue | 🟠 MajorReject unsupported Azure output formats instead of reporting them as generated.
mapFormat()only supports four formats, but the result object still reportsoptions.formatat Line 196. If a caller asks for one of the newly added formats, Azure will produce mp3/opus while the SDK labels the file as the requested format.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/AzureTTS.ts` around lines 164 - 196, The code in AzureTTS.ts constructs outputFormat via mapFormat(...) but still returns the original requested options.format in the TTSResult, which can incorrectly label audio when the requested format is unsupported; update the send-flow in the synthesize method (where outputFormat is computed and the result object is created) to validate the requested format: call mapFormat(options.format) and if it returns a different actual output (or null/undefined for unsupported) either throw a clear error rejecting unsupported formats or set result.format to the actual produced format derived from outputFormat (e.g., map outputFormat back to a canonical extension) so the returned TTSResult.format reflects the real payload; reference mapFormat, outputFormat, and the result: TTSResult creation to locate and fix the logic.src/lib/adapters/stt/azureSTTHandler.ts-471-486 (1)
471-486:⚠️ Potential issue | 🟠 MajorUse the negotiated input format for streamed audio frames.
Every chunk is tagged as
audio/x-wavhere, even if the caller is streamingmp3,ogg, orwebm. That will make Azure decode the stream incorrectly unless WAV is the only supported input and enforced up front.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/adapters/stt/azureSTTHandler.ts` around lines 471 - 486, The code currently hardcodes Content-Type: audio/x-wav when sending audio frames and the final endMessage, which breaks non-WAV streams; update the send logic in the loop and the endMessage to use the negotiated input MIME type (e.g., a variable like contentType or negotiatedAudioFormat obtained earlier in the handler) instead of the literal 'audio/x-wav', ensuring both the headerBuffer/Message and the endMessage use that variable (refer to audioStream, ws, requestId and the send logic in azureSTTHandler.ts to locate and replace the hardcoded value).src/lib/adapters/stt/azureSTTHandler.ts-369-381 (1)
369-381:⚠️ Potential issue | 🟠 Major
isFinalis wrong when interim results are enabled.This makes every successful segment non-final whenever
interimResultsis true, but the handler never emits a corresponding final segment. Downstream voice agents won't know when an utterance is actually complete.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/adapters/stt/azureSTTHandler.ts` around lines 369 - 381, The isFinal flag is computed incorrectly as "!azureOptions.interimResults", causing interim-mode segments never to be marked final; change the logic in the TranscriptionSegment construction (in azureSTTHandler) so that when interimResults is true you set isFinal based on the recognition result (e.g., data.RecognitionStatus === "Success"), and when interimResults is false you mark segments final (true). Update the isFinal assignment (referencing azureOptions, data.RecognitionStatus, and TranscriptionSegment) accordingly so final segments are emitted when the service reports Success.test/continuous-test-suite-voice.ts-333-335 (1)
333-335:⚠️ Potential issue | 🟠 MajorThe stream-handler export checks are also unconditional passes.
recordTest(..., exists || true, !exists)means this block never fails when an expected export disappears. The suite should fail here, not silently downgrade the miss to a skip.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@test/continuous-test-suite-voice.ts` around lines 333 - 335, The export existence checks in the loop use "exists || true" which makes the test always pass; update the call to recordTest in the loop over expectedExports so the second argument is the actual boolean "exists" (not short-circuited) and the skip flag is not used to silently downgrade failures (replace the third argument so it does not mark a missing export as skipped — e.g., pass false or remove the skip flag), locating the change in the loop that references expectedExports, streamHandler, and recordTest.test/continuous-test-suite-voice.ts-48-53 (1)
48-53: 🛠️ Refactor suggestion | 🟠 MajorUse a type alias for
TestResult.
interfaceis forbidden in this repo; switch this totype TestResult = { ... }.As per coding guidelines "Never use
interface. Always usetype X = { ... }."🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@test/continuous-test-suite-voice.ts` around lines 48 - 53, Replace the forbidden interface declaration with a type alias: change the `interface TestResult { ... }` to `type TestResult = { name: string; passed: boolean; skipped?: boolean; error?: string; }`, preserving all property names and optional markers exactly and updating any imports/usages if your editor flags them; ensure the symbol `TestResult` remains exported/visible as before.test/continuous-test-suite-voice.ts-347-356 (1)
347-356:⚠️ Potential issue | 🟠 MajorImport the canonical voice types module here.
This check is pointed at
../src/lib/voice/types.js, but the PR moves shared voice types intosrc/lib/types/voice.ts. In its current form the suite will skip the real type-export check instead of exercising the canonical module.As per coding guidelines "
src/lib/types/**/*.ts: All type definitions must go insrc/lib/types/. Never create type files inside feature subdirectories."🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@test/continuous-test-suite-voice.ts` around lines 347 - 356, The test currently imports "../src/lib/voice/types.js" which is outdated; update the dynamic import to the canonical types module moved to src/lib/types by importing "../src/lib/types/voice.js" (so the runtime .js path is used), keep the existing error handling and recordTest calls (symbols: the dynamic import expression and recordTest) so the suite actually exercises the shared type-export module instead of the old feature-local path.src/cli/commands/voice.ts-359-475 (1)
359-475: 🛠️ Refactor suggestion | 🟠 MajorMove these flag definitions back into the CLI command factory.
This command group is defining its option schema inline instead of consuming the centralized CLI option definitions. That makes the new voice commands drift-prone relative to the rest of the CLI surface.
Based on learnings "All CLI command options and flag definitions are centralized in
src/cli/factories/commandFactory.ts."🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/cli/commands/voice.ts` around lines 359 - 475, The voice command group is defining option schemas inline in createVoiceCommands (affecting the "synthesize", "transcribe", and "providers" subcommands) instead of using the centralized CLI option definitions; refactor createVoiceCommands to import and reuse the shared option definitions from src/cli/factories/commandFactory.ts (e.g., the centralized voice/tts/stt option objects or helper like getOption/optionFactory) and replace the .option/.positional inline calls for flags like "provider", "voice", "output", "format", "speed", "pitch", "play", "file", "language", "diarization", "word-timestamps", and "type" so handleSynthesize, handleTranscribe, and handleProviders receive the standardized args shape; ensure you update imports and types in createVoiceCommands accordingly and remove the duplicated inline schemas.src/lib/voice/providers/OpenAITTS.ts-134-176 (1)
134-176:⚠️ Potential issue | 🟠 MajorDon't silently coerce unsupported output formats to mp3.
mapFormat()only handles four formats, butTTSResult.formatstill echoesoptions.formatat Line 172. If a caller requests one of the newly added formats, the request is sent as mp3 while the SDK reports the original format, which will produce mislabeled files and confusing playback failures.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/OpenAITTS.ts` around lines 134 - 176, The code sends the request using the mapped responseFormat but returns TTSResult.format as the original options.format, causing mislabeled outputs; update OpenAITTS.mapFormat usage so the resolved format is used in the result (e.g., set result.format to the mapped responseFormat or a validated canonical format) and ensure sample rate is derived from that canonical format via getSampleRate; alternatively validate options.format up-front in the method (using mapFormat) and throw an error for unsupported formats instead of silently coercing to "mp3". Ensure references: mapFormat, responseFormat, TTSResult, options.format, and getSampleRate are updated accordingly.test/continuous-test-suite-voice.ts-122-131 (1)
122-131:⚠️ Potential issue | 🟠 MajorAlign the TTS provider matrix with the providers this PR actually ships.
src/lib/voice/index.ts:146-197only exports Google, ElevenLabs, OpenAI, and Azure TTS providers, but this list also includes Polly/PlayHT/Deepgram/Cartesia and treats missing modules as SKIP. That makes the suite pass even when the expected provider set is wrong.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@test/continuous-test-suite-voice.ts` around lines 122 - 131, The ttsProviders matrix in test/continuous-test-suite-voice.ts includes providers not exported by src/lib/voice/index.ts, causing tests to silently SKIP extras; update the ttsProviders array (symbol: ttsProviders) to only include the providers actually exported (GoogleTTS, ElevenLabsTTS, OpenAITTS, AzureTTS) and remove PollyTTS, PlayHTTTS, DeepgramTTS, CartesiaTTS entries, ensuring the envKey values match the exported providers' expected env vars and the test no longer masks missing-module failures.src/lib/voice/providers/GoogleTTS.ts-90-92 (1)
90-92:⚠️ Potential issue | 🟠 MajorDon't report service-account auth as supported until it's implemented.
isConfigured()returns true when onlycredentialsPathis set, butgetAccessToken()always returns an empty string after logging a warning. That sends callers down a “configured” path that can only 401.🔧 Minimal safe fix
isConfigured(): boolean { - return this.apiKey !== null || this.credentialsPath !== null; + return this.apiKey !== null; }Also applies to: 127-129, 232-234, 371-375
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/GoogleTTS.ts` around lines 90 - 92, isConfigured() and similar checks currently return true when credentialsPath is set even though getAccessToken() only logs a warning and returns an empty string; change those checks (e.g., isConfigured(), anywhere using credentialsPath like the other occurrences referenced) to only report configured/supported when a usable auth method exists (apiKey is non-null) or implement real service-account token retrieval; specifically, update isConfigured() to require this.apiKey (and update the other identical checks) so callers aren't misled into a configured path that will 401, or alternatively implement getAccessToken() to actually exchange the service account credentialsPath for a valid token before leaving credentialsPath treated as supported.src/lib/adapters/stt/deepgramSTTHandler.ts-95-105 (1)
95-105:⚠️ Potential issue | 🟠 MajorKeep
SUPPORTED_FORMATSandgetContentType()in sync.
mp4,mpeg, andmpgaare advertised as supported on Lines 95-105, but Lines 538-554 fall back toaudio/wavfor all three. Requests for those formats will be mislabeled on upload.Also applies to: 538-554
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/adapters/stt/deepgramSTTHandler.ts` around lines 95 - 105, SUPPORTED_FORMATS lists "mp4", "mpeg", and "mpga" but getContentType() currently falls back to "audio/wav" for those, causing incorrect Content-Type on upload; update the getContentType() implementation to return the proper MIME types for these extensions (e.g., "mp4" -> "audio/mp4" and both "mpeg" and "mpga" -> "audio/mpeg") and ensure any fallback remains only for unknown extensions so SUPPORTED_FORMATS and getContentType() stay in sync.src/lib/voice/providers/GoogleTTS.ts-84-87 (1)
84-87:⚠️ Potential issue | 🟠 MajorUse the TTS-specific env var here.
This constructor reads
GOOGLE_API_KEY, but this provider is the Google Cloud TTS integration. A setup that only definesGOOGLE_TTS_API_KEYwill look unconfigured unless callers pass the key explicitly.🔧 Proposed fix
- this.apiKey = apiKey ?? process.env.GOOGLE_API_KEY ?? null; + this.apiKey = apiKey ?? process.env.GOOGLE_TTS_API_KEY ?? null;Based on learnings
In the NeuroLink TTS SDK (src/lib/tts/), use GOOGLE_TTS_API_KEY environment variable specifically for Google Cloud Text-to-Speech access. GOOGLE_AI_API_KEY does not provide TTS access and should not be used for TTS functionality.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/GoogleTTS.ts` around lines 84 - 87, The constructor in GoogleTTS.ts sets this.apiKey from GOOGLE_API_KEY which is incorrect for the Google Cloud TTS provider; update the precedence in the GoogleTTS constructor so it uses the explicit parameter first, then process.env.GOOGLE_TTS_API_KEY, and then null (do not use GOOGLE_AI_API_KEY or GOOGLE_API_KEY for TTS), while leaving this.credentialsPath to continue using process.env.GOOGLE_APPLICATION_CREDENTIALS as before.src/lib/voice/providers/ElevenLabsTTS.ts-68-130 (1)
68-130:⚠️ Potential issue | 🟠 MajorHonor
languageCodeingetVoices(), or drop the parameter.
getVoices("ja")still returns the full voice list. Right now the argument only bypasses the cache, and every mapped voice gets the same hardcodedlanguageCodes, so filtered discovery is inaccurate.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/ElevenLabsTTS.ts` around lines 68 - 130, The getVoices(languageCode?) implementation currently ignores the requested language and hardcodes languageCodes and languageCode values; update getVoices to honor the languageCode param by deriving each voice's supported languages from the ElevenLabs response (use fields on ElevenLabsVoicesResponse.voice such as language(s)/labels if present), set voice.languageCode dynamically (e.g., primary lang) and voice.languageCodes from the actual voice metadata, then filter the returned voices to only those that include the requested languageCode; also adjust caching so either cache per-language (keyed by languageCode) or cache the full list and apply the language filter at return time, and continue to use mapGender(…) and voice.voice_id/voice.name as before.src/lib/voice/providers/OpenAISTT.ts-194-196 (1)
194-196:⚠️ Potential issue | 🟠 MajorNormalize locale tags before sending
languageto OpenAI's Whisper API.The
languageparameter must use ISO-639-1 format (2-letter codes likeen). OpenAI's transcription API does not accept locale tags likeen-USorpt-BR. Strip to the base language code before appending.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/OpenAISTT.ts` around lines 194 - 196, The code currently appends options.language directly to the multipart form in OpenAISTT (where formData.append("language", options.language) is called); normalize the locale by extracting the base ISO-639-1 code (take substring before '-' or '_' and lowercase) and validate it is a 2-letter alpha code before appending. Update the logic in the OpenAISTT method that builds formData to transform options.language -> baseLang = options.language.split(/[-_]/)[0].toLowerCase() and only call formData.append("language", baseLang) if baseLang matches /^[a-z]{2}$/.src/lib/adapters/stt/googleSTTHandler.ts-184-189 (1)
184-189:⚠️ Potential issue | 🟠 MajorCapability claims "streaming" but
transcribeStreamis not implemented.
getCapabilities()returns["stt", "streaming"], but this class has notranscribeStreammethod. Consumers relying on capability checks will incorrectly assume streaming is available.🔧 Option 1: Remove streaming capability
getCapabilities(): VoiceCapability[] { - return ["stt", "streaming"]; + return ["stt"]; }🔧 Option 2: Add placeholder streaming method
async *transcribeStream( _audioStream: AsyncIterable<Buffer>, _options: STTOptions, ): AsyncIterable<TranscriptionSegment> { throw new Error("Streaming not implemented for google-stt adapter"); }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/adapters/stt/googleSTTHandler.ts` around lines 184 - 189, getCapabilities claims "streaming" but there is no transcribeStream implemented; add an explicit placeholder async generator method transcribeStream in the googleSTT handler class (signature: async *transcribeStream(_audioStream: AsyncIterable<Buffer>, _options: STTOptions): AsyncIterable<TranscriptionSegment>) that immediately throws a clear Error("Streaming not implemented for google-stt adapter") so consumers see that streaming is unavailable, or alternatively remove "streaming" from getCapabilities() if you prefer to signal no streaming support; reference getCapabilities and transcribeStream when making the change.src/lib/voice/providers/GoogleSTT.ts-482-492 (1)
482-492:⚠️ Potential issue | 🟠 MajorPlaceholder
getAccessTokenwill cause silent auth failures.When
credentialsPathis set butapiKeyis not, the code path at lines 271-273 callsgetAccessToken(), which returns an empty string. This will sendAuthorization: Bearerheaders, causing 401 errors. Either implement service account auth or throw an error indicating it's unsupported.🔧 Proposed fix to throw instead of returning empty
private async getAccessToken(): Promise<string> { - // In production, this would use the Google Auth library - // For now, return empty (API key should be used instead) - logger.warn( - "[GoogleSTTHandler] Service account auth not implemented, use API key", - ); - return ""; + throw new Error( + "Service account authentication not implemented. Please use GOOGLE_API_KEY instead.", + ); }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/GoogleSTT.ts` around lines 482 - 492, The current getAccessToken() in GoogleSTT returns an empty string which causes silent 401s when credentialsPath is provided but apiKey is not; update getAccessToken() (in class GoogleSTT/GoogleSTTHandler) to throw a descriptive Error (e.g., "Service account auth not implemented; provide apiKey or implement service account flow") instead of returning "" so the caller fails fast, and ensure any callers of getAccessToken() will surface that error (no silent Authorization: Bearer headers).src/lib/adapters/stt/gladiaSTTHandler.ts-316-322 (1)
316-322:⚠️ Potential issue | 🟠 MajorWebSocket with headers won't work in Node.js without dynamic import.
The code uses the global
WebSocketconstructor with a headers option. In Node.js, there is no globalWebSocket, and even if polyfilled, the standard WebSocket API doesn't support custom headers. Other streaming handlers in this PR (e.g.,DeepgramSTT) dynamically import thewspackage. This will throw aReferenceErrorin Node.js.🔧 Proposed fix using dynamic import
+ // Import WebSocket for Node.js + const { default: WebSocket } = await import("ws"); + // Create WebSocket connection const ws = new WebSocket(wsUrl, { - // `@ts-expect-error` - headers are supported by Node.js WebSocket libraries headers: { "x-gladia-key": this.apiKey, }, });🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/adapters/stt/gladiaSTTHandler.ts` around lines 316 - 322, The WebSocket instantiation using the global WebSocket with a headers option will fail in Node.js; update the ws creation in gladiaSTTHandler (the block that creates const ws = new WebSocket(wsUrl, { headers: { "x-gladia-key": this.apiKey } })) to dynamically import the 'ws' package at runtime (e.g., const { WebSocket: NodeWebSocket } = await import('ws')) and then instantiate NodeWebSocket with wsUrl and the headers option so Node supports custom headers; ensure any types/ts-expect-error are removed or adjusted and that wsUrl and this.apiKey are passed to the imported WebSocket constructor.src/lib/voice/providers/OpenAIRealtime.ts-437-456 (1)
437-456:⚠️ Potential issue | 🟠 MajorMissing
call_idfrom the event breaks function-call response.The
response.function_call_arguments.doneevent includes acall_idfield that must be passed back in thefunction_call_outputitem. Using the functionnameinstead (line 508) will cause the OpenAI Realtime API to reject or misroute the response.🔧 Proposed fix to capture and use `call_id`
case "response.function_call_arguments.done": { const funcEvent = event as { name?: string; arguments?: string; + call_id?: string; }; - if (funcEvent.name && funcEvent.arguments) { + if (funcEvent.name && funcEvent.arguments && funcEvent.call_id) { try { const args = JSON.parse(funcEvent.arguments) as Record< string, unknown >; - this.handleFunctionCall(funcEvent.name, args); + this.handleFunctionCall(funcEvent.name, args, funcEvent.call_id); } catch {And update
handleFunctionCall:private async handleFunctionCall( name: string, args: Record<string, unknown>, + callId: string, ): Promise<void> { try { const result = await this.emitFunctionCall(name, args); // Send function result back if (this.ws && this.isConnected()) { this.ws.send( JSON.stringify({ type: "conversation.item.create", item: { type: "function_call_output", - call_id: name, // Note: This should be the actual call_id from the event + call_id: callId, output: JSON.stringify(result), }, }),🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/OpenAIRealtime.ts` around lines 437 - 456, The handler for the "response.function_call_arguments.done" event currently reads only funcEvent.name and funcEvent.arguments and calls handleFunctionCall(name, args); update it to also read funcEvent.call_id (or callId) from the incoming event and pass that call_id through so the subsequent function_call_output uses the original call_id rather than the function name; specifically, in the case block for "response.function_call_arguments.done" extract call_id from the event object (alongside name and arguments), parse arguments into args as before, and call handleFunctionCall(funcEvent.name, args, funcEvent.call_id) or otherwise forward the call_id to the code that constructs the function_call_output item so the outgoing item includes call_id.src/lib/neurolink.ts-11805-11813 (1)
11805-11813:⚠️ Potential issue | 🟠 MajorThese convenience APIs bypass the configured voice providers.
If the caller omits
provider, these methods hard-codeelevenlabs/deepgraminstead of honoring the provider already configured viainitializeVoice(). That makes streaming and voice discovery drift from the SDK's active voice configuration.Also applies to: 11905-11908, 12091-12094
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/neurolink.ts` around lines 11805 - 11813, The streaming convenience APIs (e.g., synthesizeStream) hard-code a fallback provider ("elevenlabs"/"deepgram") when caller omits provider, which bypasses the voice provider configured via initializeVoice(); update these calls to default to the SDK's active provider instead of a string literal by calling the VoiceFactory method that returns the currently configured provider (e.g., use VoiceFactory.getActiveProvider() or VoiceFactory.getConfiguredProvider() as available) and pass that into VoiceFactory.createTTSProvider/createSTTProvider; make the same change for the other affected streaming methods that currently use provider ?? "elevenlabs" or provider ?? "deepgram" so they respect the initialized voice/STT configuration.src/lib/neurolink.ts-11746-11753 (1)
11746-11753:⚠️ Potential issue | 🟠 MajorForward the full
CompositeVoiceConfig.
initializeVoice()acceptsCompositeVoiceConfig, but this constructor call dropsttsOptions/sttOptionsplusstreamingandlatencyMode. Those caller-supplied settings are silently ignored.🧩 Proposed fix
this.compositeVoice = new CompositeVoice({ ttsProvider: config.ttsProvider, sttProvider: config.sttProvider, defaultTTSOptions: config.defaultTTSOptions, defaultSTTOptions: config.defaultSTTOptions, + ttsOptions: config.ttsOptions, + sttOptions: config.sttOptions, trackHistory: config.trackHistory ?? true, maxHistoryTurns: config.maxHistoryTurns ?? 50, + streaming: config.streaming, + latencyMode: config.latencyMode, });🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/neurolink.ts` around lines 11746 - 11753, The CompositeVoice is being constructed with only a subset of fields, dropping caller-supplied ttsOptions, sttOptions, streaming and latencyMode from the CompositeVoiceConfig; update the constructor call that creates new CompositeVoice(...) so it forwards the entire CompositeVoiceConfig (e.g., pass config or explicitly include ttsOptions, sttOptions, streaming, latencyMode along with ttsProvider, sttProvider, defaultTTSOptions, defaultSTTOptions, trackHistory and maxHistoryTurns) to preserve caller settings and keep any defaults/nullable handling the same in initializeVoice/CompositeVoice.src/lib/neurolink.ts-11931-11950 (1)
11931-11950:⚠️ Potential issue | 🟠 MajorFallback transcription can yield an empty stream.
When the provider lacks
transcribeStream(), this branch only re-emitsresult.segments. Providers that return plainresult.textwithout segment metadata will produce no items, so the transcription is lost.🎙️ Proposed fix
- if (result.segments) { + if (result.segments?.length) { for (const segment of result.segments) { yield { ...segment, start: segment.start ?? segment.startTime ?? 0, end: segment.end ?? segment.endTime ?? 0, confidence: segment.confidence ?? 0, } as TranscriptionSegment; } + } else if (result.text) { + yield { + text: result.text, + start: 0, + end: 0, + confidence: result.confidence ?? 0, + isFinal: true, + } as TranscriptionSegment; }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/neurolink.ts` around lines 11931 - 11950, The fallback branch that collects audio and calls sttProvider.transcribe only yields result.segments, which drops transcriptions when the provider returns only result.text; update the branch in neurolink.ts (the block using audioStream, sttProvider.transcribe, result.segments and TranscriptionSegment) to handle non-segment results by yielding a single TranscriptionSegment when segments are absent: construct a segment with text from result.text (or result.transcript), start = 0, end = 0 (or best-effort duration if available), and confidence = result.confidence ?? 0, preserving existing behavior for providers that do return result.segments.src/lib/neurolink.ts-11741-11745 (1)
11741-11745:⚠️ Potential issue | 🟠 MajorDon't log raw voice config.
CompositeVoiceConfigcan carry provider config objects, so loggingconfighere can leak credentials into debug logs. Guard the serialization and sanitize it first.As per coding guidelines: use `logger.shouldLog("debug")` to guard expensive serialization before logging and `transformParamsForLogging()` to safely strip secrets before logging.🔒 Proposed fix
- logger.debug("[NeuroLink] Initializing voice capabilities", { config }); + if (logger.shouldLog("debug")) { + logger.debug("[NeuroLink] Initializing voice capabilities", { + config: transformParamsForLogging( + config as unknown as Record<string, unknown>, + ), + }); + }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/neurolink.ts` around lines 11741 - 11745, The debug log in initializeVoice currently logs the raw CompositeVoiceConfig which may contain secrets; change it to first check logger.shouldLog("debug") and only then call transformParamsForLogging(config) and log the sanitized result (e.g., logger.debug("[NeuroLink] Initializing voice capabilities", { config: transformParamsForLogging(config) })), ensuring you reference the initializeVoice method, CompositeVoiceConfig, logger.shouldLog("debug") and transformParamsForLogging() when updating the logging call.src/lib/neurolink.ts-12497-12567 (1)
12497-12567:⚠️ Potential issue | 🟠 Major
validateVoiceProvider()rejects supported STT providers.This classifier never routes
google-sttorazure-sttthroughVoiceFactory.createSTTProvider(), so those providers currently fall through toUnknown provider.🛠️ Proposed fix
- } else if ( - ["deepgram", "gladia", "whisper", "assemblyai"].includes(provider) - ) { + } else if ( + [ + "deepgram", + "gladia", + "whisper", + "assemblyai", + "google-stt", + "azure-stt", + ].includes(provider) + ) { const sttProvider = await VoiceFactory.createSTTProvider( - provider as "deepgram" | "gladia" | "whisper", + provider as + | "deepgram" + | "gladia" + | "whisper" + | "assemblyai" + | "google-stt" + | "azure-stt", );🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/neurolink.ts` around lines 12497 - 12567, validateVoiceProvider currently treats "google-stt" and "azure-stt" as unknown because the STT branch only checks ["deepgram","gladia","whisper","assemblyai"]; update that branch to include "google-stt" and "azure-stt" in the array and in the TypeScript cast for VoiceFactory.createSTTProvider (e.g., add "google-stt" | "azure-stt" to the union), so createSTTProvider(...) will be called for those providers and their validateConfig path is executed.src/lib/voice/providers/GeminiLive.ts-446-476 (1)
446-476:⚠️ Potential issue | 🟠 MajorFunction call errors are logged but not propagated to caller.
In
handleFunctionCall(), errors are caught and logged (Lines 471-474) but the function result is not sent back to Gemini on failure. The model may hang waiting for a response.Suggested fix - send error response to Gemini
} catch (err: unknown) { logger.error( `[GeminiLiveHandler] Function call failed: ${err instanceof Error ? err.message : String(err)}`, ); + // Send error response to Gemini so it doesn't hang + if (this.ws && this.isConnected()) { + const errorResponse = { + toolResponse: { + functionResponses: [ + { + id: callId, + name, + response: { error: err instanceof Error ? err.message : String(err) }, + }, + ], + }, + }; + this.ws.send(JSON.stringify(errorResponse)); + this.pendingFunctionCalls.delete(callId); + } }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/GeminiLive.ts` around lines 446 - 476, handleFunctionCall currently logs errors but never replies to Gemini, which can leave the model waiting; in the catch block for handleFunctionCall, send a toolResponse over this.ws (after verifying this.ws && this.isConnected()) that mirrors the successful response shape but contains an error payload (e.g., { id: callId, name, error: { message: errMessage, type: errType? } }), JSON.stringify it, and call this.pendingFunctionCalls.delete(callId) so the pending call is cleared; retain the existing logger.error call and ensure you guard the send behind the same connection check used elsewhere (this.ws && this.isConnected()) and include the callId and name in the error response so Gemini can correlate it.
🟡 Minor comments (8)
test/continuous-test-suite-voice.ts-84-90 (1)
84-90:⚠️ Potential issue | 🟡 MinorThis realtime-provider assertion never fails.
realtimeProviders.length >= 0is always true, so this check won't catch a broken registry. Use> 0if the intent is to verify at least one realtime provider is wired.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@test/continuous-test-suite-voice.ts` around lines 84 - 90, The test asserts realtimeProviders.length >= 0 which is always true; change the assertion in the test using factory1.getProvidersByType("realtime") and recordTest to require at least one provider by using realtimeProviders.length > 0 (i.e., update the call that currently does recordTest("Realtime providers available", realtimeProviders.length >= 0) to use > 0 so the check will fail when no realtime providers are registered).src/lib/adapters/stt/deepgramSTTHandler.ts-318-323 (1)
318-323:⚠️ Potential issue | 🟡 MinorThis appears to be dead code—the active implementation is in
src/lib/voice/providers/DeepgramSTT.tswhich correctly uses thewsnpm package.The technical concern is valid: the built-in WHATWG
WebSocketconstructor does not support a{ headers: ... }options object as shown (the second argument is reserved for subprotocols only). However, this file (deepgramSTTHandler.ts) is not imported or used anywhere in the codebase. The actual streaming implementation inDeepgramSTT.tsusesimport("ws")which does support headers, so authentication works correctly there.If this file is intended to be kept as an alternative implementation, update it to use the
wspackage or an alternative auth method (query parameters, cookies, or post-connection auth). Otherwise, remove it to avoid confusion.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/adapters/stt/deepgramSTTHandler.ts` around lines 318 - 323, This file contains dead/incorrect code using the WHATWG WebSocket constructor with a headers option (const ws = new WebSocket(wsUrl, { headers: { Authorization: `Token ${this.apiKey}` } })), which doesn’t support headers; either remove the unused deepgramSTTHandler.ts entirely to avoid confusion, or update the implementation to use the ws npm package like in DeepgramSTT.ts (import or dynamic import of "ws") and construct the socket using new Ws(wsUrl, { headers: { Authorization: `Token ${this.apiKey}` } }) and adjust types accordingly so authentication works as intended; ensure you update any exports/usage so the codebase references the correct implementation.src/lib/voice/providers/GoogleSTT.ts-88-91 (1)
88-91:⚠️ Potential issue | 🟡 Minor
supportsStreaming = trueis misleading for a chunked-batch implementation.The
transcribeStreammethod buffers ~5 seconds of audio and calls the synchronoustranscribeAPI repeatedly. This isn't true streaming (it has 5s latency per segment and no interim results). Consider settingsupportsStreaming = falseor documenting this limitation clearly.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/GoogleSTT.ts` around lines 88 - 91, The supportsStreaming flag on GoogleSTT is misleading because transcribeStream buffers ~5s and calls the synchronous transcribe method (no true low-latency streaming or interim results); update the implementation by either setting the public readonly supportsStreaming property to false on GoogleSTT or explicitly document the limitation in the class/method comments and README, and add a note in the transcribeStream method header explaining it performs chunked-batch buffering (~5s latency) and does not provide interim results; reference the supportsStreaming property and transcribeStream method when making the change.src/lib/voice/providers/DeepgramSTT.ts-526-530 (1)
526-530:⚠️ Potential issue | 🟡 MinorMissing connection timeout for WebSocket.
The
await new Promisefor connection open has no timeout. If the WebSocket fails to connect (network issues, wrong URL), the method hangs indefinitely. Other handlers in this PR use timeouts (e.g.,GladiaSTTHandleruses 10s).🔧 Proposed fix with timeout
// Wait for connection await new Promise<void>((resolve, reject) => { + const timeout = setTimeout(() => { + ws.close(); + reject(STTError.streamError("WebSocket connection timeout", "deepgram")); + }, 10000); + - ws.on("open", () => resolve()); - ws.on("error", reject); + ws.on("open", () => { + clearTimeout(timeout); + resolve(); + }); + ws.on("error", (err) => { + clearTimeout(timeout); + reject(err); + }); });🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/DeepgramSTT.ts` around lines 526 - 530, The connection wait for the WebSocket (ws) currently awaits forever; add a timeout like other handlers (e.g., GladiaSTTHandler's 10s) so the Promise rejects if "open" doesn't occur in time. Modify the Promise that listens to ws.on("open") and ws.on("error") to also set a timer (clear it on open/error) which rejects after 10_000ms, and ensure any created timer is cleaned up to avoid leaks; update the calling method in DeepgramSTT (the block that awaits new Promise for ws) to handle the rejection accordingly.src/lib/voice/providers/OpenAIRealtime.ts-448-448 (1)
448-448:⚠️ Potential issue | 🟡 MinorAsync function call handler invoked without
await.
handleFunctionCallisasyncbut is called synchronously in the switch case. Any errors will become unhandled promise rejections instead of being caught and emitted viaemitError.🛠️ Proposed fix
- this.handleFunctionCall(funcEvent.name, args); + void this.handleFunctionCall(funcEvent.name, args).catch((err) => { + this.emitError(err instanceof Error ? err : new Error(String(err))); + });🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/OpenAIRealtime.ts` at line 448, The switch case is invoking the async method handleFunctionCall(funcEvent.name, args) without awaiting it, causing unhandled promise rejections; modify the caller (the switch handler) to either await this.handleFunctionCall(...) or explicitly attach a catch that forwards errors to this.emitError(err) — ensure the enclosing function is marked async if you add await, or use this.handleFunctionCall(...).catch(err => this.emitError(err)) so all errors are caught and emitted.src/lib/types/voice.ts-900-912 (1)
900-912:⚠️ Potential issue | 🟡 MinorType guard incorrectly requires optional property.
isTranscriptionSegmentchecks thatobj.index === "number"(Line 908), butTranscriptionSegment.indexis optional (index?: numberat Line 201). Valid segments without an index will fail this guard.Suggested fix
export function isTranscriptionSegment( value: unknown, ): value is TranscriptionSegment { if (!value || typeof value !== "object") { return false; } const obj = value as Record<string, unknown>; return ( - typeof obj.index === "number" && + (obj.index === undefined || typeof obj.index === "number") && typeof obj.text === "string" && typeof obj.isFinal === "boolean" ); }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/types/voice.ts` around lines 900 - 912, isTranscriptionSegment incorrectly requires the optional property index to be present; update the guard in isTranscriptionSegment so it accepts objects where obj.index is either undefined or a number (e.g., replace the strict typeof obj.index === "number" check with a conditional that allows undefined or a number), keeping the existing checks for obj.text and obj.isFinal so the function still narrows to TranscriptionSegment when appropriate.src/lib/voice/voiceAgent.ts-68-69 (1)
68-69:⚠️ Potential issue | 🟡 MinorReadonly property is mutated.
this.configis declaredreadonlyat Line 68, butupdateVoiceSettings()mutatesthis.config.voiceSettingsat Lines 469-472. Either remove thereadonlymodifier or use a separate mutable settings field.Proposed fix
Option 1 - Remove readonly (if mutation is intended):
- private readonly config: VoiceAgentConfig; + private config: VoiceAgentConfig;Option 2 - Use a separate mutable field:
private readonly config: VoiceAgentConfig; + private voiceSettingsOverrides: Partial<VoiceAgentConfig["voiceSettings"]> = {};Also applies to: 465-475
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/voiceAgent.ts` around lines 68 - 69, The property this.config (type VoiceAgentConfig) is declared readonly but updateVoiceSettings mutates this.config.voiceSettings; fix by either removing the readonly modifier from the config declaration so mutations are allowed, or keep config readonly and introduce a new mutable field (e.g., mutableVoiceSettings or voiceSettings) that updateVoiceSettings reads/writes instead; update all references in updateVoiceSettings and any other methods that modify or rely on voiceSettings to use the chosen mutable field, leaving the original this.config as an immutable configuration object.src/lib/voice/voiceFactory.ts-554-576 (1)
554-576:⚠️ Potential issue | 🟡 MinorStatic
has*Providermethods may return false before initialization completes.The
hasTTSProvider,hasSTTProvider, andhasRealtimeProvidermethods callgetType()which reads fromtypeMapsynchronously. If called before initialization completes, they will incorrectly returnfalse.Suggested fix - add async variants or document limitation
/** * Check if a TTS provider exists (static convenience method) + * NOTE: Returns false if factory is not yet initialized. + * Use ensureInitialized() first for reliable results. */ static hasTTSProvider(nameOrAlias: string): boolean {Or add async variants:
static async hasTTSProviderAsync(nameOrAlias: string): Promise<boolean> { const factory = VoiceFactory.getInstance(); await factory.ensureInitialized(); return factory.getType(nameOrAlias) === "tts"; }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/voiceFactory.ts` around lines 554 - 576, The static hasTTSProvider/hasSTTProvider/hasRealtimeProvider methods call getType() synchronously and can return false if initialization hasn't completed; to fix, add async variants (e.g., hasTTSProviderAsync, hasSTTProviderAsync, hasRealtimeProviderAsync) that call VoiceFactory.getInstance().ensureInitialized() before using getType(), or alternatively make the existing static methods await ensureInitialized() and return Promise<boolean>; reference the existing methods hasTTSProvider/hasSTTProvider/hasRealtimeProvider, the instance method getType, and the initialization helper ensureInitialized when implementing the change so callers get correct results after initialization.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 6cb9af04-3973-4fe4-82dc-719ede7a4c8c
📒 Files selected for processing (33)
src/cli/commands/voice.tssrc/cli/parser.tssrc/lib/adapters/stt/assemblyaiSTTHandler.tssrc/lib/adapters/stt/azureSTTHandler.tssrc/lib/adapters/stt/deepgramSTTHandler.tssrc/lib/adapters/stt/gladiaSTTHandler.tssrc/lib/adapters/stt/googleSTTHandler.tssrc/lib/adapters/stt/whisperSTTHandler.tssrc/lib/neurolink.tssrc/lib/types/index.tssrc/lib/types/tts.tssrc/lib/types/voice.tssrc/lib/voice/RealtimeVoiceAPI.tssrc/lib/voice/STTProvider.tssrc/lib/voice/audio-utils.tssrc/lib/voice/compositeVoice.tssrc/lib/voice/errors.tssrc/lib/voice/index.tssrc/lib/voice/providers/AzureSTT.tssrc/lib/voice/providers/AzureTTS.tssrc/lib/voice/providers/DeepgramSTT.tssrc/lib/voice/providers/ElevenLabsTTS.tssrc/lib/voice/providers/GeminiLive.tssrc/lib/voice/providers/GoogleSTT.tssrc/lib/voice/providers/GoogleTTS.tssrc/lib/voice/providers/OpenAIRealtime.tssrc/lib/voice/providers/OpenAISTT.tssrc/lib/voice/providers/OpenAITTS.tssrc/lib/voice/stream-handler.tssrc/lib/voice/voiceAgent.tssrc/lib/voice/voiceFactory.tssrc/lib/voice/voiceRegistry.tstest/continuous-test-suite-voice.ts
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
4d38652 to
5cc6346
Compare
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
5cc6346 to
f4cd22b
Compare
|
@coderabbitai review |
✅ Actions performedReview triggered.
|
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
|
@coderabbitai full review |
✅ Actions performedFull review triggered. |
There was a problem hiding this comment.
Actionable comments posted: 9
Note
Due to the large number of review comments, Critical severity comments were prioritized as inline comments.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
src/cli/factories/commandFactory.ts (1)
318-323:⚠️ Potential issue | 🟡 MinorKeep CLI TTS formats in sync with
AudioFormat.The public type now accepts
m4a,flac,webm,mp4,mpeg, andmpga, but the CLI still rejects them and caststtsFormatto the old four-format union.🐛 Proposed fix
+ AudioFormat, AnthropicAuthConfig, AnthropicAuthMethod, ClaudeSubscriptionTier, } from "../../lib/types/index.js";ttsFormat: { type: "string" as const, - choices: ["mp3", "wav", "ogg", "opus"], + choices: ["mp3", "wav", "ogg", "opus", "m4a", "flac", "webm", "mp4", "mpeg", "mpga"], default: "mp3", description: "Audio output format", },- ttsFormat: argv.ttsFormat as "mp3" | "wav" | "ogg" | "opus" | undefined, + ttsFormat: argv.ttsFormat as AudioFormat | undefined,format: - (enhancedOptions.ttsFormat as "mp3" | "wav" | "ogg" | "opus") || - undefined, + (enhancedOptions.ttsFormat as AudioFormat | undefined) || + undefined,format: - (enhancedOptions.ttsFormat as "mp3" | "wav" | "ogg" | "opus") || - undefined, + (enhancedOptions.ttsFormat as AudioFormat | undefined) || + undefined,Also applies to: 727-729, 2696-2698, 2966-2968
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/cli/factories/commandFactory.ts` around lines 318 - 323, The CLI option ttsFormat is restricted to the old four values and must be updated to match the public AudioFormat union; locate the ttsFormat option in commandFactory.ts (and the other ttsFormat occurrences) and replace the hard-coded choices with the full set of AudioFormat values (include "mp3","wav","ogg","opus","m4a","flac","webm","mp4","mpeg","mpga") or, better, derive choices from the shared AudioFormat type/enum so they stay in sync; keep the existing default (e.g., "mp3") and ensure any casting or type annotation accepts the expanded union.src/lib/types/tts.ts (1)
153-159:⚠️ Potential issue | 🟠 MajorUpdate runtime audio format validation for the expanded
AudioFormatunion.
AudioFormatnow allowsm4a,flac,webm,mp4,mpeg, andmpga, butVALID_AUDIO_FORMATSstill rejects them viaisValidTTSOptions()andisTTSResult().🐛 Proposed fix
export const VALID_AUDIO_FORMATS: readonly AudioFormat[] = [ "mp3", "wav", "ogg", "opus", + "m4a", + "flac", + "webm", + "mp4", + "mpeg", + "mpga", ];🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/types/tts.ts` around lines 153 - 159, VALID_AUDIO_FORMATS is out of sync with the expanded AudioFormat union, causing isValidTTSOptions and isTTSResult to reject valid formats; update VALID_AUDIO_FORMATS to include the new formats (m4a, flac, webm, mp4, mpeg, mpga) so runtime validation matches the AudioFormat type, and ensure any logic in isValidTTSOptions() and isTTSResult() references VALID_AUDIO_FORMATS rather than hardcoded lists so future enum additions stay consistent.
♻️ Duplicate comments (1)
src/lib/voice/providers/AzureSTT.ts (1)
38-41:⚠️ Potential issue | 🟠 Major
supportsStreamingstill overstates Azure STT behavior.
transcribeStream()buffers ~5 seconds and repeatedly calls the batch REST endpoint, so callers do not get true continuous streaming/interim-result semantics. Either set this tofalseor implement Azure Speech SDK/WebSocket streaming.Does Azure Speech REST endpoint /speech/recognition/conversation/cognitiveservices/v1 provide true continuous streaming interim transcription results, or only final REST recognition responses?Also applies to: 282-336
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/AzureSTT.ts` around lines 38 - 41, The supportsStreaming flag on AzureSTT incorrectly claims true streaming; update the AzureSTT implementation by setting the public readonly supportsStreaming property to false (or alternatively implement true streaming via the Azure Speech SDK/WebSocket) and ensure transcribeStream()'s behavior and documentation reflect batched REST calls (buffering ~5s and returning final results) rather than interim streaming; reference the AzureSTT supportsStreaming property and the transcribeStream() method and adjust any callers/tests that rely on streaming semantics accordingly.
🟡 Minor comments (10)
src/lib/adapters/stt/deepgramSTTHandler.ts-250-256 (1)
250-256:⚠️ Potential issue | 🟡 MinorHardcoded
encoding/sample_rateignores caller-provided audio format.
transcribeStreamforce-setsencoding=linear16andsample_rate=16000regardless of what the caller is actually sending throughaudioStream. If the audio source is 48 kHz Opus (browser MediaRecorder), 44.1 kHz PCM, or anything else, Deepgram will mis-transcribe or drop characters without any error surfacing. Read these fromoptions(mirroring the batch path'sgetContentType) and fall back to sensible defaults only when unspecified.- queryParams.set("encoding", "linear16"); - queryParams.set("sample_rate", "16000"); + queryParams.set( + "encoding", + (deepgramOptions as { encoding?: string }).encoding ?? "linear16", + ); + queryParams.set( + "sample_rate", + String((deepgramOptions as { sampleRate?: number }).sampleRate ?? 16000), + ); queryParams.set("interim_results", "true");🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/adapters/stt/deepgramSTTHandler.ts` around lines 250 - 256, transcribeStream currently overrides caller audio format by hardcoding encoding and sample_rate; modify transcribeStream to read encoding and sample_rate from the provided DeepgramSTTOptions (deepgramOptions) or derive them the same way as the batch path (use getContentType or the options fields) before calling buildQueryParams, and only set defaults (e.g., linear16/16000) when those option values are absent; update the code around buildQueryParams/deepgramOptions to conditionally set queryParams.set("encoding", ...) and queryParams.set("sample_rate", ...) from the resolved values so the WS URL reflects the actual audio format being sent.src/lib/adapters/stt/deepgramSTTHandler.ts-338-386 (1)
338-386:⚠️ Potential issue | 🟡 MinorReplace
STREAMING_NOT_SUPPORTEDwithSTREAM_ERRORfor connection timeouts and mid-stream errors.The code at lines 339 and 380 uses
STT_ERROR_CODES.STREAMING_NOT_SUPPORTEDfor network/timeout failures, but this code semantically means "the provider doesn't support streaming." Callers branching on error codes (e.g., falling back to batch mode) will be misled into permanently disabling streaming for transient failures. UseSTREAM_ERRORinstead—it's the appropriate code for streaming connection and runtime failures.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/adapters/stt/deepgramSTTHandler.ts` around lines 338 - 386, Replace the incorrect error code usage: change STT_ERROR_CODES.STREAMING_NOT_SUPPORTED to STT_ERROR_CODES.STREAM_ERROR for the STTError thrown in the WebSocket timeout handler (the timeout callback that constructs new STTError) and for the STTError thrown when state.errorMessage is set in the receive loop; keep the rest of the STTError fields (category, severity, retriable, provider "deepgram") unchanged so callers see a transient STREAM_ERROR instead of a permanent STREAMING_NOT_SUPPORTED.src/lib/types/voice.ts-440-499 (1)
440-499:⚠️ Potential issue | 🟡 MinorAdd metadata for all valid
AudioFormatvalues.
AudioFormatincludesmp4,mpeg, andmpga, butAUDIO_FORMAT_DETAILShas no entries for them. Any lookup by a valid format will getundefined.🐛 Proposed fix
webm: { format: "webm", mimeType: "audio/webm", extension: ".webm", supportsStreaming: true, sampleRates: [44100, 48000], bitDepths: [16], }, + mp4: { + format: "mp4", + mimeType: "audio/mp4", + extension: ".mp4", + supportsStreaming: false, + sampleRates: [44100, 48000], + bitDepths: [16], + }, + mpeg: { + format: "mpeg", + mimeType: "audio/mpeg", + extension: ".mpeg", + supportsStreaming: true, + sampleRates: [8000, 16000, 22050, 24000, 44100, 48000], + bitDepths: [16], + }, + mpga: { + format: "mpga", + mimeType: "audio/mpeg", + extension: ".mpga", + supportsStreaming: true, + sampleRates: [8000, 16000, 22050, 24000, 44100, 48000], + bitDepths: [16], + }, };🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/types/voice.ts` around lines 440 - 499, AUDIO_FORMAT_DETAILS is missing entries for valid AudioFormat values (mp4, mpeg, mpga), causing lookups to return undefined; update the AUDIO_FORMAT_DETAILS object to include metadata objects for "mp4", "mpeg", and "mpga" (matching the shape of AudioFormatDetails: format, mimeType, extension, supportsStreaming, sampleRates, bitDepths) so every AudioFormat key is present — e.g., add mp4 (format: "mp4", mimeType: "audio/mp4", extension: ".mp4", supportsStreaming: false, sampleRates: [...], bitDepths: [...]), mpeg (format: "mpeg", mimeType: "audio/mpeg", extension: ".mpeg" or ".mpg", supportsStreaming: true, sampleRates: [...], bitDepths: [...]), and mpga (format: "mpga", mimeType: "audio/mpeg", extension: ".mpga", supportsStreaming: true, sampleRates: [...], bitDepths: [...]) ensuring sampleRates/bitDepths choices align with similar entries (e.g., mp3/opus) and types remain Partial<Record<AudioFormat, AudioFormatDetails>>.src/lib/neurolink.ts-11888-11968 (1)
11888-11968:⚠️ Potential issue | 🟡 MinorKeep
createEvaluationPipelineattached to its JSDoc.The new voice methods now sit between the evaluation pipeline docblock and
createEvaluationPipeline, so generated API docs may orphan that documentation. Move the voice methods above the “Evaluation & Scoring API” section or move thecreateEvaluationPipelineJSDoc down to line 11968.📝 Proposed structure
+ // ======================================== + // Voice API + // ======================================== + /** * Synthesize text to speech. */ async synthesize(...) { ... } async transcribe(...) { ... } async startRealtimeVoice(...) { ... } /** * Create an evaluation pipeline with the specified configuration or preset. * Pipelines orchestrate multiple scorers to evaluate AI responses comprehensively. */ async createEvaluationPipeline(...)🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/neurolink.ts` around lines 11888 - 11968, The JSDoc for createEvaluationPipeline has been separated from its function by the new voice methods; to fix, ensure the createEvaluationPipeline JSDoc stays immediately above the createEvaluationPipeline declaration by moving either the voice methods (TTS synthesize, STT transcribe, startRealtimeVoice / RealtimeProcessor.connect) above the "Evaluation & Scoring API" section or by relocating the createEvaluationPipeline JSDoc block down so it directly precedes the createEvaluationPipeline function; update references to startRealtimeVoice, synthesize, transcribe as needed to preserve logical grouping and documentation generation.src/lib/voice/providers/AzureTTS.ts-226-228 (1)
226-228:⚠️ Potential issue | 🟡 MinorEscape SSML attribute values, not just text nodes.
voiceis inserted into SSML attributes without XML attribute escaping. A malformed or user-provided voice value can break the generated SSML.🛡️ Proposed direction
- return azureOptions.ssmlTemplate - .replace("{text}", this.escapeXml(text)) - .replace("{voice}", voice); + return azureOptions.ssmlTemplate + .replace("{text}", this.escapeXml(text)) + .replace("{voice}", this.escapeXml(voice)); ... - <voice name="${voice}"> + <voice name="${this.escapeXml(voice)}">Also applies to: 245-248
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/AzureTTS.ts` around lines 226 - 228, The SSML assembly inserts the voice identifier into attributes without XML-attribute escaping, which can break SSML if voice contains quotes or special chars; update the code that builds SSML (the expression using azureOptions.ssmlTemplate.replace("{text}", this.escapeXml(text)).replace("{voice}", voice) and the analogous replacements around 245-248) to escape attribute values before insertion (e.g., call a new or existing escapeXmlAttribute/escapeXmlForAttribute helper on voice and any other attribute substitutions) so only attribute-safe strings are injected into the template.src/lib/adapters/stt/assemblyaiSTTHandler.ts-405-405 (1)
405-405:⚠️ Potential issue | 🟡 MinorRemove the forbidden non-null assertions flagged by static analysis.
These are already under guard conditions; assign guarded locals instead of using
!so the code passes the quality gate.♻️ Example cleanup
- yield segments.shift()!; + const nextSegment = segments.shift(); + if (nextSegment) { + yield nextSegment; + }- Authorization: this.apiKey!, + Authorization: this.apiKey,For private methods, pass
apiKeyas an argument or add a local guard before building headers.Also applies to: 432-432, 571-571, 604-604, 686-687
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/adapters/stt/assemblyaiSTTHandler.ts` at line 405, Replace all non-null assertion usages (e.g., the `segments.shift()!` call and other `!` usages at the referenced locations) by assigning the potentially-null value to a guarded local variable and performing an explicit null check before usage; for example, `const segment = segments.shift(); if (!segment) { /* handle or continue */ } yield segment;`. Do the same for other occurrences (lines noted around 432, 571, 604, 686–687): assign guarded locals and early-return/throw/continue as appropriate instead of using `!`. For private methods that build headers and currently rely on a non-null `apiKey`, either accept `apiKey` as a parameter or add a local guard (e.g., `const key = this.apiKey; if (!key) throw new Error(...)`) before constructing the headers, and use the guarded `key` variable. Ensure all replacements reference the original symbols (`segments`, `segment`, `apiKey`, header-building helpers) so static analysis no longer sees non-null assertions.src/lib/voice/providers/ElevenLabsTTS.ts-82-110 (1)
82-110:⚠️ Potential issue | 🟡 MinorApply the requested language filter in
getVoices().When
languageCodeis provided, this still returns every voice while markinglanguageCode: "en"for each. Consumers asking for"fr"or"de"will receive unfiltered English-labeled results.🐛 Proposed fix
- const voices: TTSVoice[] = data.voices.map((voice) => ({ + let voices: TTSVoice[] = data.voices.map((voice) => ({ ... - // Cache voices + if (languageCode) { + const requested = languageCode.toLowerCase(); + voices = voices.filter((voice) => + voice.languageCodes?.some((code) => + code.toLowerCase().startsWith(requested), + ), + ); + } + + // Cache voices🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/ElevenLabsTTS.ts` around lines 82 - 110, getVoices() currently ignores the languageCode parameter and hardcodes languageCode: "en" and a static languageCodes list; fix by filtering data.voices to only include voices that support the requested languageCode (e.g., check voice.language_codes, voice.languages, or voice.labels for supported languages) and set each TTSVoice.languageCode to the requested languageCode (or to the voice's primary supported language if languageCode is undefined). Update the voices mapping (referencing voice.voice_id, name, mapGender) to derive languageCodes from the voice metadata instead of the hardcoded array, apply the filter before creating the voices array, and keep caching behavior (voicesCache) unchanged so caching only happens when no languageCode is passed.src/lib/voice/RealtimeVoiceAPI.ts-155-183 (1)
155-183:⚠️ Potential issue | 🟡 MinorDetach realtime event handlers when connect fails.
handler.on(handlers)runs beforehandler.connect(). If connection fails, those callbacks remain attached and may fire during a later session.♻️ Proposed fix
} catch (err: unknown) { + if (handlers) { + handler.off(); + } + if (err instanceof RealtimeError) { throw err; }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/RealtimeVoiceAPI.ts` around lines 155 - 183, The handlers currently registered via handler.on(handlers) remain attached if handler.connect(mergedConfig) throws; update the connect flow so you unregister those handlers on failure: after calling handler.on(handlers) and before awaiting handler.connect, ensure the catch block calls a corresponding detach (e.g., handler.off(handlers) or handler.removeListener/removeAllListeners as appropriate) to remove the previously attached callbacks, then rethrow the RealtimeError (preserving the existing RealtimeError.connectionFailed construction); reference the existing handler.on, handler.connect and handlers symbols and add the detach call in the catch path for non-RealtimeError failures.src/lib/voice/index.ts-107-122 (1)
107-122:⚠️ Potential issue | 🟡 MinorExport the STT providers added outside
voice/providers.
AssemblyAISTTHandleris included in this PR undersrc/lib/adapters/stt/assemblyaiSTTHandler.ts, but it is not exposed from the consolidated voice entrypoint like the other STT providers.♻️ Proposed export
+export { AssemblyAISTTHandler } from "../adapters/stt/assemblyaiSTTHandler.js";🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/index.ts` around lines 107 - 122, The AssemblyAISTT provider added in src/lib/adapters/stt/assemblyaiSTTHandler.ts is not exported from the consolidated voice entrypoint; add an export in src/lib/voice/index.ts following the existing pattern (export the class and an alias) so the symbol AssemblyAISTT and AssemblyAISTTHandler (or the actual exported class name from assemblyaiSTTHandler.ts) are exposed alongside AzureSTT, DeepgramSTT, GoogleSTT, and OpenAISTT; ensure the export path points to the adapters/stt/assemblyaiSTTHandler module and matches the class name exported there.src/lib/types/realtime.ts-60-68 (1)
60-68:⚠️ Potential issue | 🟡 MinorJSDoc comments are misaligned with their fields.
The JSDoc lines
/** Turn detection mode */and/** Instructions/system prompt for the session */both land abovetemperature?, andtemperature?'s own/** Temperature for AI responses */sits aboveinstructions?. Tools like TSDoc / IDE tooltips will attribute each doc block to the wrong field.🔧 Reorder to match fields
- /** VAD threshold (0-1) */ - vadThreshold?: number; - /** Turn detection mode */ - /** Instructions/system prompt for the session */ - /** Temperature for AI responses */ - temperature?: number; - instructions?: string; - turnDetection?: "server_vad" | "manual"; + /** VAD threshold (0-1) */ + vadThreshold?: number; + /** Temperature for AI responses */ + temperature?: number; + /** Instructions/system prompt for the session */ + instructions?: string; + /** Turn detection mode */ + turnDetection?: "server_vad" | "manual";🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/types/realtime.ts` around lines 60 - 68, The JSDoc comments above the RealtimeSession type fields are misaligned so IDEs will show wrong tooltips; move each comment to immediately precede its corresponding field (`temperature?: number;` should have `/** Temperature for AI responses */`, `instructions?: string;` should have `/** Instructions/system prompt for the session */`, `turnDetection?: "server_vad" | "manual";` should have `/** Turn detection mode */`, and `tools?: RealtimeTool[];` should have `/** Tools/functions available to the model */`) so the comments attach to the correct symbols (`temperature`, `instructions`, `turnDetection`, `tools`, and type `RealtimeTool`).
🧹 Nitpick comments (7)
src/lib/voice/providers/OpenAITTS.ts (1)
101-108:getVoiceslanguage filter is a no-op.Both branches return
OpenAITTS.VOICESunchanged, solanguageCodehas no effect. Either drop the parameter/branch, or actually filter (e.g. return[]or a warning for unsupported languages). The current shape misleads callers into thinking language filtering happens.♻️ Proposed simplification
- async getVoices(languageCode?: string): Promise<TTSVoice[]> { - // OpenAI voices are pre-defined, filter by language if provided - if (languageCode && !languageCode.startsWith("en")) { - // OpenAI TTS works with multiple languages but voices are English-named - return OpenAITTS.VOICES; - } + async getVoices(_languageCode?: string): Promise<TTSVoice[]> { + // OpenAI voices are multilingual; the language parameter does not affect + // voice selection, so the full list is always returned. return OpenAITTS.VOICES; }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/OpenAITTS.ts` around lines 101 - 108, The getVoices method in OpenAITTS (function getVoices) currently ignores languageCode because both branches return OpenAITTS.VOICES; fix by either removing the languageCode parameter/branch entirely or implement real filtering: inspect OpenAITTS.VOICES for a language or locale property and return only matching entries, or if voices truly only support English return an empty array (or a logged warning) when languageCode is present and doesn't start with "en" (e.g. if languageCode && !languageCode.startsWith("en") return []). Update the logic in getVoices and adjust any callers if you remove the parameter.src/cli/factories/sagemakerCommandFactory.ts (1)
242-259:tempClientis constructed but never exercised — truthy check is dead.
new SageMakerClient(...)either throws or returns a non-null object, so!tempClientis alwaysfalse. The block validates nothing beyond what the subsequent!secureConfig.accessKeyId/!secureConfig.secretAccessKeychecks already cover. Either drop the client instantiation or actually perform a lightweight call (e.g.ListEndpointsCommandwith a small page) to validate credentials, since the comment claims "this will throw if credentials are invalid" but no network call is made.♻️ Proposed simplification
private static validateSecureConfiguration( secureConfig: SecureConfiguration, ): void { - // Create temporary AWS SDK client with secure credentials - const tempClient = new SageMakerClient({ - region: secureConfig.region, - credentials: { - accessKeyId: secureConfig.accessKeyId, - secretAccessKey: secureConfig.secretAccessKey, - }, - }); - - // Test basic connectivity (this will throw if credentials are invalid) - // Note: We're not actually making a call here, just validating the client can be created - if ( - !tempClient || - !secureConfig.accessKeyId || - !secureConfig.secretAccessKey - ) { + if (!secureConfig.accessKeyId || !secureConfig.secretAccessKey) { throw new Error("Invalid AWS credentials provided"); }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/cli/factories/sagemakerCommandFactory.ts` around lines 242 - 259, The tempClient instantiation (SageMakerClient) is not actually validating credentials because the truthy check on tempClient is useless; replace the dead check by performing a lightweight API call to validate credentials (or remove the client creation if you prefer no runtime validation). Concretely, after creating tempClient call await tempClient.send(new ListEndpointsCommand({MaxResults: 1})) (or another cheap call), catch any thrown error and convert it into the existing throw new Error("Invalid AWS credentials provided") (including the original error message for debugging); ensure you import ListEndpointsCommand, await the promise, and only treat missing secureConfig.accessKeyId/secretAccessKey as a pre-check before calling the API.src/lib/server/voice/voiceWebSocketHandler.ts (1)
241-290: Consider guarding against runaway reconnects and timer re-entrancy.Two small resilience gaps in the Soniox plumbing:
connectSoniox()retries every 500ms with no backoff or attempt cap. If Soniox is down or returns a permanent auth/quota error, each session will hammer the endpoint indefinitely and spam logs until the client disconnects.startKeepAlive()overwriteskeepAliveTimerwithout clearing an existing timer. Today the only call site is insideopen, which is paired withstopKeepAlive()onclose, so it's safe; but a stray double-invocation would leak the previous interval. A defensive clear is essentially free.🔧 Proposed tweak
- function startKeepAlive() { - keepAliveTimer = setInterval(() => { + function startKeepAlive() { + stopKeepAlive(); + keepAliveTimer = setInterval(() => { if (sonioxWs?.readyState === WebSocket.OPEN) { sonioxWs.send(JSON.stringify({ type: "keepalive" })); } }, 8000); }For backoff, consider tracking a
reconnectAttemptscounter per session and usingMath.min(500 * 2 ** attempts, 30000)with a reset on successfulopen.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/server/voice/voiceWebSocketHandler.ts` around lines 241 - 290, Add a per-session reconnectAttempts counter used by connectSoniox() to stop runaway reconnects and apply exponential backoff: on ws.close (when !sessionClosed) schedule reconnect using delay = Math.min(500 * 2**reconnectAttempts, 30000), increment reconnectAttempts each attempt, and reset reconnectAttempts = 0 in the ws.on("open") handler; also ensure you check sessionClosed before scheduling. Defensively handle keepAliveTimer in startKeepAlive()/stopKeepAlive(): clear any existing keepAliveTimer at the start of startKeepAlive() before creating a new setInterval and keep the stopKeepAlive() behavior to clear and null it. Use the existing symbols connectSoniox, startKeepAlive, stopKeepAlive, keepAliveTimer, sonioxWs, and sessionClosed to locate and apply these changes.src/lib/core/baseProvider.ts (1)
1037-1054:!ttsProviderbranch is unreachable.
ttsProvider = options.tts?.provider ?? options.provider ?? this.providerName.this.providerNameis anAIProviderNameset in the constructor and is always a non-empty string, so!ttsProvidercan never be true and the warn/return path at Line 1040–1054 is dead. If the intent was to only warn when a caller-specified provider was missing, the check should fire before falling back tothis.providerName; otherwise the branch can be dropped.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/core/baseProvider.ts` around lines 1037 - 1054, The guard that logs and returns when !ttsProvider is unreachable because ttsProvider is computed as options.tts?.provider ?? options.provider ?? this.providerName (and this.providerName is always set); update the code by either (A) moving the missing-provider check to run before falling back to this.providerName (e.g., inspect options.tts?.provider and options.provider directly and log/return if both are absent when the caller expected a specific provider), or (B) remove the !ttsProvider branch entirely and only keep the aiResponse emptiness check; locate the ttsProvider assignment and the subsequent if (!aiResponse || !ttsProvider) block in baseProvider.ts and apply one of these fixes so the warning path is reachable or eliminated.src/lib/voice/RealtimeVoiceAPI.ts (1)
380-393: Make realtime cleanup await disconnects.
clearHandlers()starts disconnects and immediately clears state, so tests or callers can proceed while sockets are still closing.♻️ Proposed direction
- static clearHandlers(): void { + static async clearHandlers(): Promise<void> { ... - handler.disconnect().catch(() => { + await handler.disconnect().catch(() => { // Ignore errors during cleanup });🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/RealtimeVoiceAPI.ts` around lines 380 - 393, clearHandlers() starts async disconnects but clears this.handlers and this.sessions immediately; change it to await all handler.disconnect() promises before clearing state. Iterate this.sessions (or this.handlers) to collect promises from handler.disconnect() (guard with handler?.isConnected()), use Promise.allSettled to wait for completion and ignore individual errors, then call this.handlers.clear(), this.sessions.clear(), and logger.debug only after all disconnects have settled. Ensure you reference the existing clearHandlers(), this.sessions, this.handlers, and handler.disconnect() symbols when making the change.src/lib/voice/providers/DeepgramSTT.ts (1)
491-492: Floating promise fromsendAudio().
sendAudio()is invoked withoutvoidor a.catch(). Any rejection inside (other than the caught block at Line 484) would become an unhandled rejection. TheOpenAIRealtime/Azure/Gladia equivalents all usevoid sendAudio(); please match.- // Start sending audio in background - sendAudio(); + // Start sending audio in background + void sendAudio();Also,
(error as Error).messageat Line 497 is an unnecessary cast —erroris already typedError | nulland the enclosingif (error)has already narrowed it.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/DeepgramSTT.ts` around lines 491 - 492, Change the floating promise by calling sendAudio() with the void operator (void sendAudio()) so any rejection is intentionally ignored the same way as other providers; also remove the unnecessary cast "(error as Error).message" and use "error.message" directly since error is already typed and narrowed in the enclosing if block. Ensure changes are applied around the sendAudio invocation and the error handling in the same function (sendAudio / surrounding scope).src/lib/voice/providers/OpenAIRealtime.ts (1)
80-108: Connect-phase listeners leak; later errors can cross-fire into a settled promise.The
open/errorhandlers registered viathis.ws!.on(...)at Lines 85 and 90 are never removed after the connection promise settles. Afterconnect()succeeds, a subsequent socket error will still invoke the oldreject(err)(no-op on a settled promise) alongside the intended handler at Line 106, and the orphanopenhandler will remain forever. Useonce(or remove them in afinally) so that only the persistentmessage/close/errorhandlers installed after connect remain.♻️ Suggested refactor
- this.ws!.on("open", () => { - clearTimeout(timeout); - resolve(); - }); - - this.ws!.on("error", (err) => { - clearTimeout(timeout); - reject(err); - }); + this.ws!.once("open", () => { + clearTimeout(timeout); + resolve(); + }); + this.ws!.once("error", (err) => { + clearTimeout(timeout); + reject(err); + });Same pattern applies to
waitForSessionCreated(Lines 292–322) — useonceor explicitoffin both resolve and reject paths (current code only removes on matching events, so timeouts leak the listener).🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/voice/providers/OpenAIRealtime.ts` around lines 80 - 108, The connect() method registers persistent "open" and "error" handlers with this.ws using .on which are never removed, causing cross-fire into the settled connection promise; change those connect-phase handlers to .once (or register them and explicitly remove them in a finally block) so the promise's resolve/reject handlers are not leaked, ensure the timeout is always cleared, and keep the persistent handlers (this.ws.on("message"/"close"/"error")) only after the promise resolves; apply the same fix in waitForSessionCreated (replace ephemeral .on listeners with .once or remove them on both resolve and reject) so timeouts or later socket events do not leak listeners or trigger callbacks on a settled promise.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 22646063-aeed-476c-bc94-0951b34755ea
⛔ Files ignored due to path filters (1)
pnpm-lock.yamlis excluded by!**/pnpm-lock.yaml
📒 Files selected for processing (46)
CHANGELOG.mdmemory-bank/voice-bridge-implementation-plan.mdpackage.jsonsrc/cli/factories/commandFactory.tssrc/cli/factories/sagemakerCommandFactory.tssrc/cli/loop/optionsSchema.tssrc/lib/adapters/stt/assemblyaiSTTHandler.tssrc/lib/adapters/stt/azureSTTHandler.tssrc/lib/adapters/stt/deepgramSTTHandler.tssrc/lib/adapters/stt/gladiaSTTHandler.tssrc/lib/adapters/stt/googleSTTHandler.tssrc/lib/adapters/stt/whisperSTTHandler.tssrc/lib/core/baseProvider.tssrc/lib/factories/providerRegistry.tssrc/lib/neurolink.tssrc/lib/observability/exporters/laminarExporter.tssrc/lib/observability/exporters/posthogExporter.tssrc/lib/observability/utils/spanSerializer.tssrc/lib/server/voice/voiceWebSocketHandler.tssrc/lib/types/generate.tssrc/lib/types/index.tssrc/lib/types/realtime.tssrc/lib/types/server.tssrc/lib/types/span.tssrc/lib/types/stream.tssrc/lib/types/stt.tssrc/lib/types/tts.tssrc/lib/types/voice.tssrc/lib/utils/sttProcessor.tssrc/lib/voice/RealtimeVoiceAPI.tssrc/lib/voice/STTProvider.tssrc/lib/voice/audio-utils.tssrc/lib/voice/errors.tssrc/lib/voice/index.tssrc/lib/voice/providers/AzureSTT.tssrc/lib/voice/providers/AzureTTS.tssrc/lib/voice/providers/DeepgramSTT.tssrc/lib/voice/providers/ElevenLabsTTS.tssrc/lib/voice/providers/GeminiLive.tssrc/lib/voice/providers/GoogleSTT.tssrc/lib/voice/providers/GoogleTTS.tssrc/lib/voice/providers/OpenAIRealtime.tssrc/lib/voice/providers/OpenAISTT.tssrc/lib/voice/providers/OpenAITTS.tssrc/lib/voice/stream-handler.tstest/continuous-test-suite-voice.ts
💤 Files with no reviewable changes (2)
- CHANGELOG.md
- src/lib/types/server.ts
| { | ||
| "name": "@juspay/neurolink", | ||
| "version": "9.55.4", | ||
| "version": "9.55.2", |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Confirm what versions are already published on npm
curl -s https://registry.npmjs.org/@juspay/neurolink | jq '.["dist-tags"], ([.versions | keys[]] | sort | .[-5:])'Repository: juspay/neurolink
Length of output: 143
Version number regressed (9.55.4 → 9.55.2) — blocks npm publish.
The PR downgrades the version in package.json, but npm registry already has 9.55.4 published as the latest. Publishing this version will fail because npm rejects versions equal to or lower than what is already on the registry.
🔧 Fix
- "version": "9.55.2",
+ "version": "9.55.4",Or bump to the next appropriate version (e.g., 9.56.0 if this contains significant changes).
📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| "version": "9.55.2", | |
| "version": "9.55.4", |
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@package.json` at line 3, The package.json "version" field was regressed from
9.55.4 to 9.55.2 which will block npm publish; update the "version" value in
package.json to a forward-incremented semver (e.g., 9.56.0 or the next
appropriate patch/minor depending on changes) so it is greater than 9.55.4, and
ensure any release notes or CI/version-related checks that reference the
"version" field are consistent with this new value.
| // Create WebSocket connection | ||
| const ws = new WebSocket(wsUrl, { | ||
| // @ts-expect-error - headers are supported by Node.js WebSocket libraries | ||
| headers: { | ||
| "Ocp-Apim-Subscription-Key": this.subscriptionKey, | ||
| }, | ||
| }); | ||
|
|
There was a problem hiding this comment.
🧩 Analysis chain
🌐 Web query:
Does Node.js 18/20 built-in global WebSocket (undici) support a headers option in the constructor, or does it silently ignore non-standard options like the ws npm package allows?
💡 Result:
No, the built-in global WebSocket from undici in Node.js 18/20 does not support a headers option directly in the constructor. The constructor accepts new WebSocket(url[, protocols]), where the second argument can be a string, array of strings for subprotocols, or a WebSocketInit object primarily for dispatcher (e.g., custom agent) and protocols. Official Undici documentation and examples do not list headers as a supported option in WebSocketInit. A comment in the undici PR adding WebSocketInit (#2088) mentions it "conveniently lets people set headers for the ws handshake," but this appears Node.js-specific or outdated, as current docs (v6.x+) do not confirm headers support and examples only show dispatcher/protocols. Node.js docs and tutorials show only basic usage without options. Unlike the ws npm package, which explicitly supports { headers: {...} } and silently handles non-standard options, undici's spec-compliant WebSocket does not; passing unsupported options like headers may be ignored or error, but is not documented/supported. To set custom headers, use a custom Dispatcher (e.g., Client with headers), but this requires undici import, not global WebSocket.
Citations:
- 1: websocket: add websocketinit nodejs/undici#2088
- 2: https://undici-docs.vramana.dev/docs/api/WebSocket
- 3: https://undici-docs.vramana.dev/docs/5.x/api/WebSocket
- 4: WebSocket (and probably fetch) does not support Authorization header from URL nodejs/undici#4744
- 5: WebSockets nodejs/undici#1811
- 6: https://developer.mozilla.org/en-US/docs/Web/API/WebSocket/WebSocket
- 7: https://nodejs.dev/en/learn/getting-started/websocket
- 8: https://nodejs.org/learn/getting-started/websocket
🏁 Script executed:
# Verify azureSTTHandler.ts at lines 310-317
if [ -f "src/lib/adapters/stt/azureSTTHandler.ts" ]; then
echo "=== azureSTTHandler.ts lines 310-317 ==="
sed -n '310,317p' "src/lib/adapters/stt/azureSTTHandler.ts"
echo ""
fi
# Check DeepgramSTT.ts for the claimed correct pattern at lines 392-397
if [ -f "src/lib/adapters/stt/deepgramSTTHandler.ts" ]; then
echo "=== deepgramSTTHandler.ts lines 390-400 (context around 392-397) ==="
sed -n '390,400p' "src/lib/adapters/stt/deepgramSTTHandler.ts"
echo ""
fi
# Check gladiaSTTHandler.ts at lines 254-259
if [ -f "src/lib/adapters/stt/gladiaSTTHandler.ts" ]; then
echo "=== gladiaSTTHandler.ts lines 254-259 ==="
sed -n '254,259p' "src/lib/adapters/stt/gladiaSTTHandler.ts"
echo ""
fiRepository: juspay/neurolink
Length of output: 957
🏁 Script executed:
# Check if ws package is imported elsewhere and how event handlers are attached
echo "=== Checking for ws package imports ==="
rg -n "import.*from.*['\"]ws['\"]" src/lib/adapters/stt/
echo ""
echo "=== Checking for dynamic import of ws ==="
rg -n "await import\(.*ws" src/lib/adapters/stt/
echo ""
echo "=== Checking onmessage vs on('message') patterns ==="
rg -n "on\(['\"]message" src/lib/adapters/stt/ | head -20Repository: juspay/neurolink
Length of output: 196
🏁 Script executed:
# Get line counts to understand file sizes
wc -l src/lib/adapters/stt/azureSTTHandler.ts src/lib/adapters/stt/deepgramSTTHandler.ts src/lib/adapters/stt/gladiaSTTHandler.ts 2>/dev/null || echo "Files not found"Repository: juspay/neurolink
Length of output: 213
🏁 Script executed:
# Find where WebSocket is created in deepgramSTTHandler.ts
echo "=== Searching for WebSocket instantiation in deepgramSTTHandler.ts ==="
rg -n "new WebSocket\|await import\(.*ws\)" src/lib/adapters/stt/deepgramSTTHandler.ts -A 3 -B 1
echo ""
echo "=== Full deepgramSTTHandler.ts WebSocket-related code ==="
rg -n "WebSocket\|ws\|headers" src/lib/adapters/stt/deepgramSTTHandler.ts -A 2 -B 1 | head -40Repository: juspay/neurolink
Length of output: 189
🏁 Script executed:
# Check imports at the top of deepgramSTTHandler.ts
echo "=== deepgramSTTHandler.ts imports (first 50 lines) ==="
head -50 src/lib/adapters/stt/deepgramSTTHandler.tsRepository: juspay/neurolink
Length of output: 1437
🏁 Script executed:
# Check how event handlers are attached in both files
echo "=== Event handlers in azureSTTHandler.ts ==="
rg -n "\.on\(|onmessage|onerror|onopen|onclose" src/lib/adapters/stt/azureSTTHandler.ts | head -10
echo ""
echo "=== Event handlers in deepgramSTTHandler.ts ==="
rg -n "\.on\(|onmessage|onerror|onopen|onclose" src/lib/adapters/stt/deepgramSTTHandler.ts | head -10Repository: juspay/neurolink
Length of output: 453
🏁 Script executed:
# Find ws variable assignment in deepgramSTTHandler.ts
echo "=== Searching for ws variable in deepgramSTTHandler.ts ==="
rg -n "const ws\s*=|let ws\s*=|var ws\s*=" src/lib/adapters/stt/deepgramSTTHandler.ts -A 5
echo ""
echo "=== Broader search for WebSocket or ws in deepgramSTTHandler.ts ==="
rg -n "ws\." src/lib/adapters/stt/deepgramSTTHandler.ts | head -20Repository: juspay/neurolink
Length of output: 881
🏁 Script executed:
# Check a wider range around the event handler lines to see context
echo "=== deepgramSTTHandler.ts lines 270-290 ==="
sed -n '270,290p' src/lib/adapters/stt/deepgramSTTHandler.ts
echo ""
echo "=== deepgramSTTHandler.ts lines 300-320 ==="
sed -n '300,320p' src/lib/adapters/stt/deepgramSTTHandler.tsRepository: juspay/neurolink
Length of output: 1308
All three STT handlers (azure, deepgram, gladia) fail to send auth headers—they use undici's global WebSocket which silently ignores the headers option.
The undici WebSocket (Node 18+) conforms to WHATWG spec and does not support a headers option in the constructor. The Ocp-Apim-Subscription-Key header at lines 310–317 is never sent; Azure closes with 401. The same bug exists at:
deepgramSTTHandler.tslines 270–275 (Authorization header)gladiaSTTHandler.tslines 254–259 (x-gladia-key header)
All three require dynamic import of the ws npm package, which supports custom headers:
🔧 Fix pattern
- const ws = new WebSocket(wsUrl, {
- // `@ts-expect-error` - headers are supported by Node.js WebSocket libraries
+ const { default: NodeWebSocket } = await import("ws");
+ const ws = new NodeWebSocket(wsUrl, {
headers: {
- "Ocp-Apim-Subscription-Key": this.subscriptionKey,
+ "Ocp-Apim-Subscription-Key": this.subscriptionKey!,
},
});Event handler syntax differs: ws package uses on("message", ...) instead of onmessage property. Update all four handler assignments in each file.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/lib/adapters/stt/azureSTTHandler.ts` around lines 310 - 317, The
WebSocket constructors in azureSTTHandler (ws variable), deepgramSTTHandler, and
gladiaSTTHandler are using the WHATWG/undici global WebSocket which ignores the
headers option; replace the constructor usage by dynamically importing the "ws"
npm package (e.g., const {default: Ws} = await import("ws")) and instantiate new
Ws(wsUrl, { headers: { ... } }) so the Ocp-Apim-Subscription-Key / Authorization
/ x-gladia-key headers are actually sent; also update the event wiring from
onmessage/onopen/onclose/onerror to the "ws" package style ws.on('message',
...), ws.on('open', ...), ws.on('close', ...), ws.on('error', ...) and preserve
existing handler logic in the Azure handler (and the analogous handlers in
deepgramSTTHandler.ts and gladiaSTTHandler.ts).
| // Create WebSocket connection | ||
| const ws = new WebSocket(wsUrl, { | ||
| // @ts-expect-error - headers are supported by Node.js WebSocket libraries | ||
| headers: { | ||
| Authorization: `Token ${this.apiKey}`, | ||
| }, | ||
| }); |
There was a problem hiding this comment.
🧩 Analysis chain
🌐 Web query:
Does Deepgram streaming /v1/listen accept authentication via the token query parameter instead of the Authorization header?
💡 Result:
No, Deepgram's streaming /v1/listen endpoint (WebSocket at wss://api.deepgram.com/v1/listen) does not accept authentication via a generic "token" query parameter. Authentication must use the Authorization header with "Token YOUR_API_KEY" or "Bearer YOUR_JWT_TOKEN" (from /v1/auth/grant). For browser/client-side WebSocket connections where custom headers are blocked, use Sec-WebSocket-Protocol: "token, YOUR_API_KEY" instead. Query parameters are for options like model, language, etc., not auth tokens. Official docs confirm header-based auth exclusively, with no mention of token query param support.
Citations:
- 1: https://developers.deepgram.com/guides/fundamentals/authenticating
- 2: https://developers.deepgram.com/reference/authentication
- 3: https://developers.deepgram.com/reference/speech-to-text/listen-streaming
- 4: https://developers.deepgram.com/reference/speech-to-text/listen-streaming.mdx
- 5: https://developers.deepgram.com/docs/using-the-sec-websocket-protocol.mdx
- 6: https://developers.deepgram.com/docs/using-the-sec-websocket-protocol
- 7: https://developers.deepgram.com/reference/speech-to-text/listen-pre-recorded
- 8: https://developers.deepgram.com/guides/fundamentals/make-your-first-api-request
WebSocket auth headers won't work with WHATWG WebSocket in Node ≥22, but query parameter auth is not the solution.
This handler attempts to pass headers to the WebSocket constructor, but Node.js's built-in WHATWG WebSocket (Node ≥22) and browsers do not accept a headers option in the constructor. The second argument is reserved for subprotocols only, so the Authorization header will be silently dropped at runtime, and Deepgram will reject the connection with 401.
However, Deepgram's /v1/listen endpoint does not accept authentication via a query parameter. Authentication requires the Authorization: Token <API_KEY> header. Per Deepgram's official documentation, for client-side connections where custom headers are unavailable, use the Sec-WebSocket-Protocol header instead:
Sec-WebSocket-Protocol: token, <API_KEY>
Update the WebSocket instantiation to use the subprotocol approach for WHATWG compatibility:
🔧 Suggested fix
- const ws = new WebSocket(wsUrl, {
- // `@ts-expect-error` - headers are supported by Node.js WebSocket libraries
- headers: {
- Authorization: `Token ${this.apiKey}`,
- },
- });
+ const ws = new WebSocket(wsUrl, [`token, ${this.apiKey}`]);🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/lib/adapters/stt/deepgramSTTHandler.ts` around lines 269 - 275, The
WebSocket constructor currently passes a headers option (in the ws creation
block using wsUrl and this.apiKey) which WHATWG WebSocket (Node ≥22 / browsers)
ignores; replace that with the Sec-WebSocket-Protocol style auth by passing the
subprotocol(s) as the second argument instead of headers (e.g., supply "token,
<API_KEY>" or an array with that string) when instantiating the WebSocket (the
code that creates const ws = new WebSocket(wsUrl, { headers: { Authorization:
`Token ${this.apiKey}` } }) should be changed to call new WebSocket(wsUrl,
'token, <API_KEY>' or ['token, <API_KEY>']) using this.apiKey, and remove the
headers/ts-expect-error approach so Deepgram receives the token via
Sec-WebSocket-Protocol).
| constructor(apiKey?: string, baseUrl?: string) { | ||
| this.apiKey = apiKey ?? process.env.GLADIA_API_KEY ?? null; | ||
| this.baseUrl = baseUrl ?? "https://api.gladia.io/v2"; | ||
| } |
There was a problem hiding this comment.
🧩 Analysis chain
🌐 Web query:
Gladia STT v2 API real-time streaming WebSocket endpoint URL and initialization flow
💡 Result:
Gladia STT v2 real-time streaming uses WebSocket. There is no direct static WebSocket endpoint; instead, follow this initialization flow: 1. POST to https://api.gladia.io/v2/live with header x-gladia-key: YOUR_API_KEY and JSON body specifying audio config (required: encoding e.g. "wav/pcm", sample_rate e.g. 16000, bit_depth e.g. 16, channels e.g. 1; optional: model e.g. "solaria-1", language_config, etc.). Example curl: curl --request POST \ --url https://api.gladia.io/v2/live \ --header 'Content-Type: application/json' \ --header 'x-gladia-key: YOUR_GLADIA_API_KEY' \ --data '{ "encoding": "wav/pcm", "sample_rate": 16000, "bit_depth": 16, "channels": 1 }' Response: { "id": "uuid", "url": "wss://api.gladia.io/v2/live?token=uuid" } 2. Connect WebSocket to the returned url (wss://api.gladia.io/v2/live?token=SESSION_ID). 3. Send audio chunks as binary frames or JSON {"type": "audio_chunk", "data": {"chunk": "base64audio"}}. 4. Send {"type": "stop_recording"} to end. 5. Receive events like transcripts (type: "transcript"), speech_start, etc. To get final results: GET https://api.gladia.io/v2/live/{id}. Note: This differs from v1 which had a fixed wss://api.gladia.io/audio/text/audio-transcription with config on connect.
Citations:
- 1: https://docs.gladia.io/api-reference/v2/live/websocket
- 2: https://docs.gladia.io/chapters/live-stt/quickstart
- 3: https://gladia-95.mintlify.app/chapters/live-stt
- 4: https://docs.gladia.io/chapters/live-stt/migration-from-v1
- 5: https://gladia-95.mintlify.app/api-reference/v2/live/init
- 6: https://docs.gladia.io/api-reference/v2/live/init
- 7: https://gladia-95.mintlify.app/api-reference/v2/live/websocket
🏁 Script executed:
cat -n src/lib/adapters/stt/gladiaSTTHandler.ts | sed -n '230,250p'Repository: juspay/neurolink
Length of output: 876
🏁 Script executed:
cat -n src/lib/adapters/stt/gladiaSTTHandler.ts | sed -n '240,280p'Repository: juspay/neurolink
Length of output: 1528
🏁 Script executed:
cat -n src/lib/adapters/stt/gladiaSTTHandler.ts | sed -n '280,320p'Repository: juspay/neurolink
Length of output: 1483
🏁 Script executed:
cat -n src/lib/adapters/stt/gladiaSTTHandler.ts | sed -n '320,360p'Repository: juspay/neurolink
Length of output: 1565
WebSocket implementation uses v1 API instead of v2.
The code hardcodes the v1 WebSocket endpoint (wss://api.gladia.io/audio/text/audio-transcription, line 240) and sends configuration on connect (lines 351-359), but the constructor defaults baseUrl to https://api.gladia.io/v2 (line 116). Gladia v2 requires a different flow: POST to https://api.gladia.io/v2/live with audio configuration to obtain a session-specific token, then connect to wss://api.gladia.io/v2/live?token=SESSION_ID. The current implementation will fail against the v2 API. Update to use the v2 initialization flow or remove the v2 baseUrl if v1 is intentional.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/lib/adapters/stt/gladiaSTTHandler.ts` around lines 114 - 117, The
WebSocket code uses the v1 endpoint and sends config on connect while the
constructor defaults baseUrl to v2; update the Gladia STT adapter to follow the
v2 flow: in the constructor/initialization for GladiaSTTHandler (constructor and
the method that opens the socket), perform a POST to
https://api.gladia.io/v2/live with the audio/config payload to receive the
session token, then open the WebSocket to
wss://api.gladia.io/v2/live?token=SESSION_TOKEN and stop sending the
configuration over the socket on connect; ensure apiKey handling (this.apiKey)
is passed in the POST, and remove or change any hardcoded
wss://api.gladia.io/audio/text/audio-transcription usage so all endpoints match
the v2 flow (or if you intended v1, change the constructor default baseUrl to
the v1 base and keep the existing ws flow).
| getSupportedFormats(): AudioFormat[] { | ||
| return ["opus"]; // OpenAI Realtime uses PCM16/g711 but we expose as opus for simplicity | ||
| } |
There was a problem hiding this comment.
Emitted audio chunks are labeled as opus but actual bytes are PCM16 — consumers will fail to decode.
sendSessionUpdate configures output_audio_format: "pcm16" (Line 245) yet getSupportedFormats() advertises "opus" and every emitted RealtimeAudioChunk sets format: "opus" (Lines 340, 352). Any downstream consumer that dispatches decoding based on chunk.format will try to decode raw PCM as Opus. Either request Opus output from OpenAI or label the chunks correctly as pcm16.
🔧 Suggested fix
getSupportedFormats(): AudioFormat[] {
- return ["opus"]; // OpenAI Realtime uses PCM16/g711 but we expose as opus for simplicity
+ return ["pcm16"];
}- this.emitAudio({
- data: audioData,
- index: this.audioChunkIndex++,
- isFinal: false,
- format: "opus",
- sampleRate: 24000,
- });
+ this.emitAudio({
+ data: audioData,
+ index: this.audioChunkIndex++,
+ isFinal: false,
+ format: "pcm16",
+ sampleRate: 24000,
+ });(apply symmetrically to the response.audio.done branch)
Also applies to: 336-354
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/lib/voice/providers/OpenAIRealtime.ts` around lines 47 - 49, The
getSupportedFormats() return value and emitted RealtimeAudioChunk format labels
are incorrect: change getSupportedFormats() to include "pcm16" (or return only
"pcm16") and update every place that sets chunk.format to "opus" in
OpenAIRealtime (including the RealtimeAudioChunk construction in the audio data
handling and the response.audio.done branch) to "pcm16"; alternatively, if you
prefer Opus, modify sendSessionUpdate to request output_audio_format: "opus" and
ensure emitted chunks are actual Opus bytes—make the change consistently in
getSupportedFormats(), sendSessionUpdate, and where chunks are created to keep
format labels and bytes in sync.
f4cd22b to
d75d713
Compare
d75d713 to
604ce21
Compare
604ce21 to
579c6c7
Compare
Documentation Validation Results
🚧 Please fix the failing checks before merging. Commit: |
579c6c7 to
ad077b9
Compare
Review Feedback Addressed (Cycle 1)Files Deleted (all comments on these are now moot)
Fixes Applied to Remaining Code
STT Pipeline Bug Fixes
Validation
@coderabbitai full review |
|
Tip For best results, initiate chat on the files or code changes.
Triggering a full review now against the updated state of the branch. [full_review] 🧠 Learnings used |
ad077b9 to
ae33100
Compare
ae33100 to
636e140
Compare
|
@coderabbitai full review |
✅ Actions performedFull review triggered. |
636e140 to
4878a57
Compare
|
@coderabbitai full review |
✅ Actions performedFull review triggered. |
Add Speech-to-Text (STT) as a new capability alongside existing TTS,
with multi-provider support for both. Everything flows through the
existing generate() and stream() JSON config pattern.
New TTS providers (via generate({ tts: { provider: "..." } })):
- openai-tts: OpenAI TTS API (tts-1, tts-1-hd), 6 voices
- elevenlabs: ElevenLabs (eleven_multilingual_v2)
- azure-tts: Azure Cognitive Services Speech
New STT providers (via generate({ stt: { enabled: true, audio, provider } })):
- whisper/openai-stt: OpenAI Whisper API
- google-stt: Google Cloud Speech-to-Text
- deepgram: Deepgram Nova-2/Nova-3
- azure-stt: Azure Cognitive Services Speech
Realtime providers (registered for future SDK use):
- openai-realtime: OpenAI Realtime API (WebSocket)
- gemini-live: Google Gemini Live (WebSocket)
Infrastructure:
- STTProcessor (mirrors TTSProcessor) with SpanType.STT observability
- Audio utilities: format detection, WAV creation, PCM resampling
- ChunkedAudioStream with backpressure and validation
- 30-second fetch timeout on all voice provider API calls
- CLI flags: --stt, --stt-provider, --input-audio, --stt-language, --tts-provider
STT pipeline in generate():
- When stt.audio provided without text: transcription becomes the prompt
- When stt.audio provided with text: transcription prepended as context
- result.transcription contains STTResult with text + confidence
Tested end-to-end with real API calls:
- Google TTS, OpenAI TTS, ElevenLabs, Azure TTS (valid MP3 output)
- Whisper STT (0.95), Deepgram (1.0), Google STT (0.98), Azure STT (0.9)
- Full round-trip: Whisper→Vertex LLM→ElevenLabs (audio in → audio out)
4878a57 to
1e4be5f
Compare
|
@coderabbitai full review |
✅ Actions performedFull review triggered. |
|
Superseded by #1005 (fresh branch to fix CI workflow trigger issue) |
Summary
Consumer API
Code review fixes
Test plan
Summary by CodeRabbit
New Features
Infrastructure